Learning Click-Based Deep Structure-Preserving Embeddings with Visual Attention-Reference-Cited by-同舟云学术

Learning Click-Based Deep Structure-Preserving Embeddings with Visual Attention

Published:2019-09-18 Issue:3 Volume:15 Page:1-19
ISSN:1551-6857
Container-title:ACM Transactions on Multimedia Computing, Communications, and Applications
language:en
Short-container-title:ACM Trans. Multimedia Comput. Commun. Appl.

Author:

Li Yehao¹,Pan Yingwei²,Yao Ting²,Chao Hongyang¹,Rui Yong³,Mei Tao²

Affiliation:

1. Sun Yat-sen University, Guangzhou, China

2. JD AI Research, Beijing, China

3. Lenovo, Beijing, China

Abstract

One fundamental problem in image search is to learn the ranking functions (i.e., the similarity between query and image). Recent progress on this topic has evolved through two paradigms: the text-based model and image ranker learning. The former relies on image surrounding texts, making the similarity sensitive to the quality of textual descriptions. The latter may suffer from the robustness problem when human-labeled query-image pairs cannot represent user search intent precisely. We demonstrate in this article that the preceding two limitations can be well mitigated by learning a cross-view embedding that leverages click data. Specifically, a novel click-based Deep Structure-Preserving Embeddings with visual Attention (DSPEA) model is presented, which consists of two components: deep convolutional neural networks followed by image embedding layers for learning visual embedding, and a deep neural networks for generating query semantic embedding. Meanwhile, visual attention is incorporated at the top of the convolutional neural network to reflect the relevant regions of the image to the query. Furthermore, considering the high dimension of the query space, a new click-based representation on a query set is proposed for alleviating this sparsity problem. The whole network is end-to-end trained by optimizing a large margin objective that combines cross-view ranking constraints with in-view neighborhood structure preservation constraints. On a large-scale click-based image dataset with 11.7 million queries and 1 million images, our model is shown to be powerful for keyword-based image search with superior performance over several state-of-the-art methods and achieves, to date, the best reported NDCG@25 of 52.21%.

Funder

Guangzhou Science and Technology Program, China

National Natural Science Foundation of China

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Networks and Communications,Hardware and Architecture

Link

https://dl.acm.org/doi/pdf/10.1145/3328994

Reference46 articles.

1. Bing Bai Jason Weston David Grangier Ronan Collobert Kunihiko Sadamasa Yanjun Qi Corinna Cortes and Mehryar Mohri. 2009. Polynomial semantic indexing. In Advances in Neural Information Processing Systems. 64--72. Bing Bai Jason Weston David Grangier Ronan Collobert Kunihiko Sadamasa Yanjun Qi Corinna Cortes and Mehryar Mohri. 2009. Polynomial semantic indexing. In Advances in Neural Information Processing Systems. 64--72.

2. Bag-of-Words Based Deep Neural Network for Image Retrieval

3. Discriminative feature selection for multi-view cross-domain learning

Cited by 6 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Image sentiment considering color palette recommendations based on influence scores for image advertisement;Electronic Commerce Research;2024-05-09

2. Sparsity-guided Discriminative Feature Encoding for Robust Keypoint Detection;ACM Transactions on Multimedia Computing, Communications, and Applications;2023-12-09

3. Deep Learning-Based State-Dependent ARX Modeling and Predictive Control of Nonlinear Systems;IEEE Access;2023

4. Smart Director: An Event-Driven Directing System for Live Broadcasting;ACM Transactions on Multimedia Computing, Communications, and Applications;2021-11-30

5. A Decade Survey of Content Based Image Retrieval using Deep Learning;IEEE Transactions on Circuits and Systems for Video Technology;2021