Pith. sign in

REVIEW 3 cited by

Cross-Modal Discrete Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.05438 v1 pith:2RJ7AVND submitted 2021-06-10 cs.CV

classification cs.CV
keywords modalitiesrepresentationcross-modaldifferentembeddingacrossdiscretizedlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in representation learning have demonstrated an ability to represent information from different modalities such as video, text, and audio in a single high-level embedding vector. In this work we present a self-supervised learning framework that is able to learn a representation that captures finer levels of granularity across different modalities such as concepts or events represented by visual objects or spoken words. Our framework relies on a discretized embedding space created via vector quantization that is shared across different modalities. Beyond the shared embedding space, we propose a Cross-Modal Code Matching objective that forces the representations from different views (modalities) to have a similar distribution over the discrete embedding space such that cross-modal objects/actions localization can be performed without direct supervision. In our experiments we show that the proposed discretized multi-modal fine-grained representation (e.g., pixel/word/frame) can complement high-level summary representations (e.g., video/sentence/waveform) for improved performance on cross-modal retrieval tasks. We also observe that the discretized representation uses individual clusters to represent the same semantic concept across modalities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations

    cs.CV 2025-07 conditional novelty 6.0 of 10

    By aligning category-level information across video, audio, and flow into one unified representation while keeping modality-specific details separate, this paper shows that standard domain generalization methods impro...

  2. Open-set Cross Modal Generalization via Multimodal Unified Representation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    The authors propose OSCMG, an open-set version of Cross Modal Generalization, and show their MICU method with masked contrastive learning and unified jigsaw puzzles outperforms prior methods.

  3. MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Splitting quantization across multiple small sub-codebooks with nested masking raises VQ-VAE reconstruction fidelity, giving MGVQ rFID 0.49 and PSNR 24.70 on ImageNet at 16 times downsampling.

Pith tools