LiveBot: Generating Live Video Comments Based on Visual and Textual Contexts-Reference-Cited by-同舟云学术

LiveBot: Generating Live Video Comments Based on Visual and Textual Contexts

Published:2019-07-17 Issue: Volume:33 Page:6810-6817
ISSN:2374-3468
Container-title:Proceedings of the AAAI Conference on Artificial Intelligence
language:
Short-container-title:AAAI

Author:

Ma Shuming,Cui Lei,Dai Damai,Wei Furu,Sun Xu

Abstract

We introduce the task of automatic live commenting. Live commenting, which is also called “video barrage”, is an emerging feature on online video sites that allows real-time comments from viewers to fly across the screen like bullets or roll at the right side of the screen. The live comments are a mixture of opinions for the video and the chit chats with other comments. Automatic live commenting requires AI agents to comprehend the videos and interact with human viewers who also make the comments, so it is a good testbed of an AI agent’s ability to deal with both dynamic vision and language. In this work, we construct a large-scale live comment dataset with 2,361 videos and 895,929 live comments. Then, we introduce two neural models to generate live comments based on the visual and textual contexts, which achieve better performance than previous neural baselines such as the sequence-to-sequence model. Finally, we provide a retrieval-based evaluation protocol for automatic live commenting where the model is asked to sort a set of candidate comments based on the log-likelihood score, and evaluated on metrics such as mean-reciprocal-rank. Putting it all together, we demonstrate the first “LiveBot”. The datasets and the codes can be found at https://github.com/lancopku/livebot.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Subject

General Medicine

Cited by 24 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Generative Steganography via Live Comments on Streaming Video Frames;IEEE Transactions on Computational Social Systems;2024-06

2. Speak From Heart: An Emotion-Guided LLM-Based Multimodal Method for Emotional Dialogue Generation;Proceedings of the 2024 International Conference on Multimedia Retrieval;2024-05-30

3. Understanding Human Preferences: Towards More Personalized Video to Text Generation;Proceedings of the ACM Web Conference 2024;2024-05-13

4. Personalized time-sync comment generation based on a multimodal transformer;Multimedia Systems;2024-03-30

5. Real-time Arabic Video Captioning Using CNN and Transformer Networks Based on Parallel Implementation;Diyala Journal of Engineering Sciences;2024-03-07