Hierarchical Time-Aware Summarization with an Adaptive Transformer for Video Captioning
-
Published:2023-07-25
Issue:04
Volume:17
Page:569-592
-
ISSN:1793-351X
-
Container-title:International Journal of Semantic Computing
-
language:en
-
Short-container-title:Int. J. Semantic Computing
Author:
Cardoso Leonardo Vilela1,
Guimarães Silvio Jamil Ferzoli1,
do Patrocínio Júnior Zenilton Kleber Gonçalves1
Affiliation:
1. Image and Multimedia Data Science Laboratory (IMSCIENCE), Pontifcia Universidade Catlica de Minas Gerais (PUC Minas), Av. Dom José Gaspar, 500 - Prédio 20, 30535-901, Belo Horizonte, Brazil
Abstract
A coherent description is an ultimate goal regarding video captioning via a couple of sentences because it might also affect the consistency and intelligibility of the generated results. In this context, a paragraph describing a video is affected by the activities used to both produce its specific narrative and provide some clues that can also assist in decreasing textual repetition. This work proposes a model, named Hierarchical time-aware Summarization with an Adaptive Transformer (HSAT), that uses a strategy to enhance the frame selection reducing the amount of information that needed to be processed along with attention mechanisms to enhance a memory-augmented transformer. This new approach increases the coherence among the generated sentences, assessing data importance (about the video segments) contained in the self-attention results and uses that to improve readability using only a small fraction of time spent by the other methods. The test results show the potential of this new approach as it provides higher coherence among the various video segments, decreasing the repetition in the generated sentences and improving the description diversity in the ActivityNet Captions dataset.
Funder
Conselho Nacional de Desenvolvimento Cientifico e Tecnolóogico CNPq
Fundação de Amparo à Pesquisa do Estado de Minas Gerais FAPEMIG
Publisher
World Scientific Pub Co Pte Ltd
Subject
Artificial Intelligence,Computer Networks and Communications,Computer Science Applications,Linguistics and Language,Information Systems,Software
Cited by
1 articles.
订阅此论文施引文献
订阅此论文施引文献,注册后可以免费订阅5篇论文的施引文献,订阅后可以查看论文全部施引文献