REVIEW 3 cited by
Decoupling Common and Unique Representations for Multimodal Self-supervised Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The increasing availability of multi-sensor data sparks wide interest in multimodal self-supervised learning. However, most existing approaches learn only common representations across modalities while ignoring intra-modal training and modality-unique representations. We propose Decoupling Common and Unique Representations (DeCUR), a simple yet effective method for multimodal self-supervised learning. By distinguishing inter- and intra-modal embeddings through multimodal redundancy reduction, DeCUR can integrate complementary information across different modalities. We evaluate DeCUR in three common multimodal scenarios (radar-optical, RGB-elevation, and RGB-depth), and demonstrate its consistent improvement regardless of architectures and for both multimodal and modality-missing settings. With thorough experiments and comprehensive analysis, we hope this work can provide valuable insights and raise more interest in researching the hidden relationships of multimodal representations.
Forward citations
Cited by 3 Pith papers
-
Toward Seasonal Guidelines for Robust Deep-Learning Sentinel-2 Building Detection in Different Area Types
Summer Sentinel-2 imagery and a U-Net give the most reliable building detection; winter scenes and low-density settlement types produce large accuracy drops.
-
SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
A unified multi-modal remote sensing foundation model with adaptive patch merging, modality prompt tokens, mixture of experts, and query-based semantic aggregation contrastive learning outperforms SkySense by 1.8 poin...
-
Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis
Under limited Russian flood labels, multimodal U-Net++ (F1 0.84) outperforms fine-tuned AnySat for water mapping, and the masks feed an EMERCOM-style damage pipeline that matches Tulun 2019 official area and exposure ...
Discussion (0). Sign in to comment.