Aggregating per-token pullback metrics via the Fréchet mean on the SPD manifold outperforms Euclidean mean pooling for sentence classification, with most of the gain attributable to geometric aggregation rather than learned encoder structure.
Latent space oddity: on the curvature of deep generative models
7 Pith papers cite this work, alongside 25 external citations. Polarity classification is still indexing.
abstract
Deep generative models provide a systematic way to learn nonlinear data distributions, through a set of latent variables and a nonlinear "generator" function that maps latent points into the input space. The nonlinearity of the generator imply that the latent space gives a distorted view of the input space. Under mild conditions, we show that this distortion can be characterized by a stochastic Riemannian metric, and demonstrate that distances and interpolants are significantly improved under this metric. This in turn improves probability distributions, sampling algorithms and clustering in the latent space. Our geometric analysis further reveals that current generators provide poor variance estimates and we propose a new generator architecture with vastly improved variance estimates. Results are demonstrated on convolutional and fully connected variational autoencoders, but the formalism easily generalize to other deep generative models.
years
2026 7representative citing papers
Retrieval from out-of-domain foundation models enables personalization of a lightweight transformer for stress detection, yielding +3.92% accuracy and +4.76% F1 gains on WESAD without user labels.
Unsupervised manifold learning on ICSD data reveals a low-dimensional embedding that segregates superconductors and predicts critical temperatures across families.
LAST linearizes action manifolds with Lie-algebraic mapping and discretizes them into approximately isotropic charts to align with VL semantic geometry via Gromov-Wasserstein distance.
PACE recovers geometry-consistent continuous transport dynamics from destructive single-cell snapshots via anisotropic metrics and neural bridges, reducing reconstruction distances by 23.7% on average across seven datasets.
Generated video latents are pushed toward a shell-shaped manifold fitted to high-quality SFT video patches, producing a dense reward that reduces blur, over-smoothing, and motion artifacts in text-to-video models.
X-VAE uses empirical statistics from a pretrained autoencoder to set a data-adaptive Gaussian prior and introduces a latent scaling factor for controllable generation.
citing papers explorer
-
Riemannian Geometry for Pre-trained Language Model Embeddings
Aggregating per-token pullback metrics via the Fréchet mean on the SPD manifold outperforms Euclidean mean pooling for sentence classification, with most of the gain attributable to geometric aggregation rather than learned encoder structure.
-
Retrieval-Augmented Personalization with Foundation Models for Wearable Stress Detection
Retrieval from out-of-domain foundation models enables personalization of a lightweight transformer for stress detection, yielding +3.92% accuracy and +4.76% F1 gains on WESAD without user labels.
-
Charting the emergent low-dimensional manifold of quantum materials
Unsupervised manifold learning on ICSD data reveals a low-dimensional embedding that segregates superconductors and predicts critical temperatures across families.
-
LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment
LAST linearizes action manifolds with Lie-algebraic mapping and discretizes them into approximately isotropic charts to align with VL semantic geometry via Gromov-Wasserstein distance.
-
PACE: Geometry-Aware Bridge Transport for Single-Cell Trajectory Inference
PACE recovers geometry-consistent continuous transport dynamics from destructive single-cell snapshots via anisotropic metrics and neural bridges, reducing reconstruction distances by 23.7% on average across seven datasets.
-
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation
Generated video latents are pushed toward a shell-shaped manifold fitted to high-quality SFT video patches, producing a dense reward that reduces blur, over-smoothing, and motion artifacts in text-to-video models.
-
eXact-Prior Variational Autoencoder (X-VAE): Learning Data-Adaptive Gaussian Mixture Priors for Latent Distributions
X-VAE uses empirical statistics from a pretrained autoencoder to set a data-adaptive Gaussian prior and introduces a latent scaling factor for controllable generation.