Improvement of deep cross-modal retrieval by generating real-valued representation-Reference-Cited by-同舟云学术

Improvement of deep cross-modal retrieval by generating real-valued representation

Published:2021-04-27 Issue: Volume:7 Page:e491
ISSN:2376-5992
Container-title:PeerJ Computer Science
language:en
Short-container-title:

Author:

Bhatt Nikita¹,Ganatra Amit²^ORCID

Affiliation:

1. U & P U. Patel Department of Computer Engineering, Chandubhai S. Patel Institute of Technology, Charotar University of Science and Technology (CHARUSAT), Changa, India

2. Devang Patel Institute of Advance Technology and Research, Charotar University of Science and Technology (CHARUSAT), Changa, India

Abstract

The cross-modal retrieval (CMR) has attracted much attention in the research community due to flexible and comprehensive retrieval. The core challenge in CMR is the heterogeneity gap, which is generated due to different statistical properties of multi-modal data. The most common solution to bridge the heterogeneity gap is representation learning, which generates a common sub-space. In this work, we propose a framework called “Improvement of Deep Cross-Modal Retrieval (IDCMR)”, which generates real-valued representation. The IDCMR preserves both intra-modal and inter-modal similarity. The intra-modal similarity is preserved by selecting an appropriate training model for text and image modality. The inter-modal similarity is preserved by reducing modality-invariance loss. The mean average precision (mAP) is used as a performance measure in the CMR system. Extensive experiments are performed, and results show that IDCMR outperforms over state-of-the-art methods by a margin 4% and 2% relatively with mAP in the text to image and image to text retrieval tasks on MSCOCO and Xmedia dataset respectively.

Publisher

PeerJ

Subject

General Computer Science

Link

https://peerj.com/articles/cs-491.pdf

Reference26 articles.

1. Applications of convolutional neural networks;Bhandare;International Journal of Computer Science and Information Technologies,2016

2. Collective matrix factorization hashing for multimodal data;Ding,2014

3. Canonical correlation analysis: an overview with application to learning methods;Hardoon;Neural Computation,2004

4. Deep cross-modal hashing;Jiang,2016

5. Learning hash functions for cross-view similarity search;Kumar,2011

Cited by 4 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Exploring the Effectiveness of Binary-Valued and Real-Valued Representations for Cross-Modal Retrieval;2023-03-28

2. Medical image retrieval using a novel local relative directional edge pattern and Zernike moments;Multimedia Tools and Applications;2023-03-02

3. Challenges and New Opportunities in Diverse Approaches of Big Data Stream Analytics;Proceedings of Third International Conference on Sustainable Expert Systems;2023

4. Impact of Binary-Valued Representation on the Performance of Cross-Modal Retrieval System;International Journal of Mathematical, Engineering and Management Sciences;2022-12-01