Y^2Seq2Seq: Cross-Modal Representation Learning for 3D Shape and Text by Joint Reconstruction and Prediction of View and Word Sequences

· 2018 · cs.CV · arXiv 1811.02745

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

open full Pith review browse 1 citing papers arXiv PDF

abstract

A recent method employs 3D voxels to represent 3D shapes, but this limits the approach to low resolutions due to the computational cost caused by the cubic complexity of 3D voxels. Hence the method suffers from a lack of detailed geometry. To resolve this issue, we propose Y^2Seq2Seq, a view-based model, to learn cross-modal representations by joint reconstruction and prediction of view and word sequences. Specifically, the network architecture of Y^2Seq2Seq bridges the semantic meaning embedded in the two modalities by two coupled `Y' like sequence-to-sequence (Seq2Seq) structures. In addition, our novel hierarchical constraints further increase the discriminability of the cross-modal representations by employing more detailed discriminative information. Experimental results on cross-modal retrieval and 3D shape captioning show that Y^2Seq2Seq outperforms the state-of-the-art methods.

representative citing papers

3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale

cs.CV · 2025-11-17 · unverdicted · novelty 5.0

3DAlign-DAER improves fine-grained 3D-text alignment at scale via dynamic attention policy with Monte Carlo calibration and efficient retrieval, supported by a new Align3D-2M dataset of 2M pairs.

citing papers explorer

Showing 1 of 1 citing paper.

3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale cs.CV · 2025-11-17 · unverdicted · none · ref 1 · internal anchor
3DAlign-DAER improves fine-grained 3D-text alignment at scale via dynamic attention policy with Monte Carlo calibration and efficient retrieval, supported by a new Align3D-2M dataset of 2M pairs.

Y^2Seq2Seq: Cross-Modal Representation Learning for 3D Shape and Text by Joint Reconstruction and Prediction of View and Word Sequences

fields

years

verdicts

representative citing papers

citing papers explorer