Pith. sign in

REVIEW 2 cited by

Simultaneous Dimensionality Reduction: A Data Efficient Approach for Multimodal Representations Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.04458 v3 pith:POCAUTQS submitted 2023-10-05 stat.ML physics.data-an

classification stat.MLphysics.data-an
keywords datadimensionalitymethodsreductionanalysiscovariationlinearmuch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We explore two primary classes of approaches to dimensionality reduction (DR): Independent Dimensionality Reduction (IDR) and Simultaneous Dimensionality Reduction (SDR). In IDR methods, of which Principal Components Analysis is a paradigmatic example, each modality is compressed independently, striving to retain as much variation within each modality as possible. In contrast, in SDR, one simultaneously compresses the modalities to maximize the covariation between the reduced descriptions while paying less attention to how much individual variation is preserved. Paradigmatic examples include Partial Least Squares and Canonical Correlations Analysis. Even though these DR methods are a staple of statistics, their relative accuracy and data set size requirements are poorly understood. We introduce a generative linear model to synthesize multimodal data with known variance and covariance structures to examine these questions. We assess the accuracy of the reconstruction of the covariance structures as a function of the number of samples, signal-to-noise ratio, and the number of varying and covarying signals in the data. Using numerical experiments, we demonstrate that linear SDR methods consistently outperform linear IDR methods and yield higher-quality, more succinct reduced-dimensional representations with smaller datasets. Remarkably, regularized CCA can identify low-dimensional weak covarying structures even when the number of samples is much smaller than the dimensionality of the data, which is a regime challenging for all dimensionality reduction methods. Our work corroborates and explains previous observations in the literature that SDR can be more effective in detecting covariation patterns in data. These findings suggest that SDR should be preferred to IDR in real-world data analysis when detecting covariation is more important than preserving variation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Computational Thresholds in Multi-Modal Learning via the Spiked Matrix-Tensor Model

    stat.ML 2025-06 conditional novelty 8.0 of 10

    For a spiked matrix-tensor model with a shared latent vector, sequential matrix-then-tensor recovery reaches the optimal Bayesian weak-recovery thresholds, while joint risk minimization makes even the easy matrix part...

  2. Accurate Estimation of Mutual Information in High Dimensional Data

    physics.data-an 2025-05 conditional novelty 5.0 of 10

    Neural MI estimators can become reliable in low-latent-dimension settings with a protocol of max-test early stopping, subsampling extrapolation, and probabilistic critics.

Pith tools