Pith. sign in

REVIEW 4 cited by

Decoupling Common and Unique Representations for Multimodal Self-supervised Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.05300 v3 pith:RZJSHAL3 submitted 2023-09-11 cs.CV

classification cs.CV
keywords multimodalrepresentationscommondecurlearningself-supervisedacrossdecoupling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The increasing availability of multi-sensor data sparks wide interest in multimodal self-supervised learning. However, most existing approaches learn only common representations across modalities while ignoring intra-modal training and modality-unique representations. We propose Decoupling Common and Unique Representations (DeCUR), a simple yet effective method for multimodal self-supervised learning. By distinguishing inter- and intra-modal embeddings through multimodal redundancy reduction, DeCUR can integrate complementary information across different modalities. We evaluate DeCUR in three common multimodal scenarios (radar-optical, RGB-elevation, and RGB-depth), and demonstrate its consistent improvement regardless of architectures and for both multimodal and modality-missing settings. With thorough experiments and comprehensive analysis, we hope this work can provide valuable insights and raise more interest in researching the hidden relationships of multimodal representations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Galileo: Learning Global & Local Features of Many Remote Sensing Modalities

    cs.CV 2025-02 conditional novelty 7.0 of 10

    A single multimodal transformer, Galileo, jointly learns global and local features from optical, radar, elevation, weather, and land-cover inputs and outperforms specialized models on eleven benchmarks.

  2. Toward Seasonal Guidelines for Robust Deep-Learning Sentinel-2 Building Detection in Different Area Types

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Summer Sentinel-2 imagery and a U-Net give the most reliable building detection; winter scenes and low-density settlement types produce large accuracy drops.

  3. SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A unified multi-modal remote sensing foundation model with adaptive patch merging, modality prompt tokens, mixture of experts, and query-based semantic aggregation contrastive learning outperforms SkySense by 1.8 poin...

  4. Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Under limited Russian flood labels, multimodal U-Net++ (F1 0.84) outperforms fine-tuned AnySat for water mapping, and the masks feed an EMERCOM-style damage pipeline that matches Tulun 2019 official area and exposure ...

Pith tools