DeepStory: Video Story QA by Deep Embedded Memory Networks-Reference-Cited by-同舟云学术

DeepStory: Video Story QA by Deep Embedded Memory Networks

Published:2017-08 Issue: Volume: Page:
ISSN:
Container-title:Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence
language:
Short-container-title:

Author:

Kim Kyung-Min¹²,Heo Min-Oh¹,Choi Seong-Ho¹,Zhang Byoung-Tak¹²

Affiliation:

1. School of Computer Science and Engineering, Seoul National University

2. Surromind Robotics

Abstract

Question-answering (QA) on video contents is a significant challenge for achieving human-level intelligence as it involves both vision and language in real-world settings. Here we demonstrate the possibility of an AI agent performing video story QA by learning from a large amount of cartoon videos. We develop a video-story learning model, i.e. Deep Embedded Memory Networks (DEMN), to reconstruct stories from a joint scene-dialogue video stream using a latent embedding space of observed data. The video stories are stored in a long-term memory component. For a given question, an LSTM-based attention model uses the long-term memory to recall the best question-story-answer triplet by focusing on specific words containing key information. We trained the DEMN on a novel QA dataset of children’s cartoon video series, Pororo. The dataset contains 16,066 scene-dialogue pairs of 20.5-hour videos, 27,328 fine-grained sentences for scene description, and 8,913 story-related QA pairs. Our experimental results show that the DEMN outperforms other QA models. This is mainly due to 1) the reconstruction of video stories in a scene-dialogue combined form that utilize the latent embedding and 2) attention. DEMN also achieved state-of-the-art results on the MovieQA benchmark.

Publisher

International Joint Conferences on Artificial Intelligence Organization

Cited by 41 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Instance-Level Trojan Attacks on Visual Question Answering via Adversarial Learning in Neuron Activation Space;2024 International Joint Conference on Neural Networks (IJCNN);2024-06-30

2. Synthesizing Coherent Story with Auto-Regressive Latent Diffusion Models;2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV);2024-01-03

3. Question difficulty estimation via enhanced directional modality association transformer;Applied Intelligence;2023-10-05

4. Tem-adapter: Adapting Image-Text Pretraining for Video Question Answer;2023 IEEE/CVF International Conference on Computer Vision (ICCV);2023-10-01

5. A Video Question Answering Model Based on Knowledge Distillation;Information;2023-06-12