Double-layer affective visual question answering network-Reference-Cited by-同舟云学术

Double-layer affective visual question answering network

Published:2021 Issue:1 Volume:18 Page:155-168
ISSN:1820-0214
Container-title:Computer Science and Information Systems
language:en
Short-container-title:COMSIS J

Author:

Guo Zihan¹,Han Dezhi¹,Massetto Francisco²,Li Kuan-Ching³

Affiliation:

1. College of Information Engineering, Shanghai Maritime University Shanghai, China

2. Center for Cognition and Complex Systems, Universidade Federal do ABC (UFABC), Santo André, Brazil

3. Dept. of Computer Science and Information Engineering, Providence University Taichung, Taiwan

Abstract

Visual Question Answering (VQA) has attracted much attention recently in both natural language processing and computer vision communities, as it offers insight into the relationships between two relevant sources of information. Tremendous advances are seen in the field of VQA due to the success of deep learning. Based upon advances and improvements, the Affective Visual Question Answering Network (AVQAN) enriches the understanding and analysis of VQA models by making use of the emotional information contained in the images to produce sensitive answers, while maintaining the same level of accuracy as ordinary VQA baseline models. It is a reasonably new task to integrate the emotional information contained in the images into VQA. However, it is challenging to separate questionguided-attention from mood-guided-attention due to the concatenation of the question words and the mood labels in AVQAN. Also, it is believed that this type of concatenation is harmful to the performance of the model. To mitigate such an effect, we propose the Double-Layer Affective Visual Question Answering Network (DAVQAN) that divides the task of generating emotional answers in VQA into two simpler subtasks: the generation of non-emotional responses and the production of mood labels, and two independent layers are utilized to tackle these subtasks. Comparative experimentation conducted on a preprocessed dataset to performance comparison shows that the overall performance of DAVQAN is 7.6% higher than AVQAN, demonstrating the effectiveness of the proposed model.We also introduce more advanced word embedding method and more fine-grained image feature extractor into AVQAN and DAVQAN to further improve their performance and obtain better results than their original models, which proves that VQA integrated with affective computing can improve the performance of the whole model by improving these two modules just like the general VQA.

Publisher

National Library of Serbia

Subject

General Computer Science

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Design of knowledge incorporated VQA based on spatial GCNN with structured sentence embedding and linking algorithm;Journal of Intelligent & Fuzzy Systems;2023-12-02

2. IdenMultiSig: Identity-Based Decentralized Multi-Signature in Internet of Things;IEEE Transactions on Computational Social Systems;2023-08

3. A Visual Question Answering Network Merging High- and Low-Level Semantic Information;IEICE Transactions on Information and Systems;2023-05-01

4. A new frog leaping algorithm-oriented fully convolutional neural network for dance motion object saliency detection;Computer Science and Information Systems;2022

5. Cross-modality co-attention networks for visual question answering;Soft Computing;2021-01-05