SCENT uses VLM-generated scene descriptions as a semantic bridge to align electronic-nose signals with visual and textual embeddings, improving cross-modal smell retrieval and enabling object-context odor disentanglement.
In: ICML (2021)
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CV 2years
2026 2representative citing papers
HyFL-CLIP distills Euclidean CLIP alignment into hyperbolic space using cross-manifold similarity and Einstein midpoint aggregation to capture hierarchical part-whole relations, achieving up to 19.5% gains in long-text retrieval under perturbations.
citing papers explorer
-
What Images Cannot Say: Language-Guided Olfactory Representation Learning
SCENT uses VLM-generated scene descriptions as a semantic bridge to align electronic-nose signals with visual and textual embeddings, improving cross-modal smell retrieval and enabling object-context odor disentanglement.
-
HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding
HyFL-CLIP distills Euclidean CLIP alignment into hyperbolic space using cross-manifold similarity and Einstein midpoint aggregation to capture hierarchical part-whole relations, achieving up to 19.5% gains in long-text retrieval under perturbations.