V2Meow: Meowing to the Visual Beat via Video-to-Music Generation-Reference-Cited by-同舟云学术

V2Meow: Meowing to the Visual Beat via Video-to-Music Generation

Published:2024-03-24 Issue:5 Volume:38 Page:4952-4960
ISSN:2374-3468
Container-title:Proceedings of the AAAI Conference on Artificial Intelligence
language:
Short-container-title:AAAI

Author:

Su Kun,Li Judith Yue,Huang Qingqing,Kuzmin Dima,Lee Joonseok,Donahue Chris,Sha Fei,Jansen Aren,Wang Yu,Verzetti Mauro,Denk Timo

Abstract

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced audio codecs, the exploration of video-acoustic signatures has been confined to specific visual scenarios. In contrast, our research confronts the challenge of learning globally aligned signatures between video and music directly from paired music and videos, without explicitly modeling domain-specific rhythmic or semantic relationships. We propose V2Meow, a video-to-music generation system capable of producing high-quality music audio for a diverse range of video input types using a multi-stage autoregressive model. Trained on 5k hours of music audio clips paired with video frames mined from in-the-wild music videos, V2Meow is competitive with previous domain-specific models when evaluated in a zero-shot manner. It synthesizes high-fidelity music audio waveforms solely by conditioning on pre-trained general-purpose visual features extracted from video frames, with optional style control via text prompts. Through both qualitative and quantitative evaluations, we demonstrate that our model outperforms various existing music generation systems in terms of visual-audio correspondence and audio quality. Music samples are available at tinyurl.com/v2meow.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. GPT-4 Driven Cinematic Music Generation Through Text Processing;ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP);2024-04-14

2. Let the Beat Follow You - Creating Interactive Drum Sounds From Body Rhythm;2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV);2024-01-03