Y2Seq2Seq: Cross-Modal Representation Learning for 3D Shape and Text by Joint Reconstruction and Prediction of View and Word Sequences-Reference-Cited by-同舟云学术

Y2Seq2Seq: Cross-Modal Representation Learning for 3D Shape and Text by Joint Reconstruction and Prediction of View and Word Sequences

Published:2019-07-17 Issue: Volume:33 Page:126-133
ISSN:2374-3468
Container-title:Proceedings of the AAAI Conference on Artificial Intelligence
language:
Short-container-title:AAAI

Author:

Han Zhizhong,Shang Mingyang,Wang Xiyang,Liu Yu-Shen,Zwicker Matthias

Abstract

Jointly learning representations of 3D shapes and text is crucial to support tasks such as cross-modal retrieval or shape captioning. A recent method employs 3D voxels to represent 3D shapes, but this limits the approach to low resolutions due to the computational cost caused by the cubic complexity of 3D voxels. Hence the method suffers from a lack of detailed geometry. To resolve this issue, we propose Y2Seq2Seq, a view-based model, to learn cross-modal representations by joint reconstruction and prediction of view and word sequences. Specifically, the network architecture of Y2Seq2Seq bridges the semantic meaning embedded in the two modalities by two coupled “Y” like sequence-tosequence (Seq2Seq) structures. In addition, our novel hierarchical constraints further increase the discriminability of the cross-modal representations by employing more detailed discriminative information. Experimental results on cross-modal retrieval and 3D shape captioning show that Y2Seq2Seq outperforms the state-of-the-art methods.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Subject

General Medicine

Cited by 23 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. TriCoLo: Trimodal Contrastive Loss for Text to Shape Retrieval;2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV);2024-01-03

2. Zero3D: Semantic-Driven 3D Shape Generation for Zero-Shot Learning;Advances in Computer Graphics;2023-12-29

3. CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-Training;2023 IEEE/CVF International Conference on Computer Vision (ICCV);2023-10-01

4. From concept to space: a new perspective on AIGC-involved attribute translation;Digital Creativity;2023-07-03

5. Parts2Words: Learning Joint Embedding of Point Clouds and Texts by Bidirectional Matching Between Parts and Words;2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR);2023-06