Affiliation:
1. Content Research Division Electronics and Telecommunications Research Institute Daejeon Republic of Korea
2. Department of Electronic Engineering Inha University Incheon Republic of Korea
3. Department of Computer Engineering Yonsei University Wonju Republic of Korea
Abstract
AbstractMusic identification is widely regarded as a solved problem for music searching in quiet environments, but its performance tends to degrade in TV broadcast audio owing to the presence of dialogue or sound effects. In addition, constructing an accurate dataset for measuring the performance of background music monitoring in TV broadcast audio is challenging. We propose a framework for monitoring background music by automatic identification and introduce a background music cue sheet. The framework comprises three main components: music identification, music–speech separation, and music detection. In addition, we introduce the Cue‐K‐Drama dataset, which includes reference songs, audio tracks from 60 episodes of five Korean TV drama series, and corresponding cue sheets that provide the start and end timestamps of background music. Experimental results on the constructed and existing datasets demonstrate that the proposed framework, which incorporates music identification with music–speech separation and music detection, effectively enhances TV broadcast audio monitoring.
Funder
Ministry of Culture, Sports and Tourism
Reference33 articles.
1. G. C.Sebasti Ciurana E.Molina M.Miron O.Meyers J.Six andX.Serra BAF: an audio fingerprinting dataset for broadcast monitoring (Proc. 23rd Int. Soc. Music Inf. Retr. Conf. Bengaluru India) 2022 pp.908–916.
2. A.Wang An industrial‐strength audio search algorithm (Proc. Int. Conf. Music Inf. Retr. Baltimore USA) 2003 pp.7–13.
3. J.HaitsmaandT.Kalker A highly robust audio fingerprinting system (Proc. Int. Soc. Music Inf. Retr. Conf. Paris France) 2002 pp.107–115.
4. A large TV dataset for speech and music activity detection;Hung Y.‐N.;EURASIP J. Audio Speech Music Process.,2022
5. Open Broadcast Media Audio from TV: A Dataset of TV Broadcast Audio with Relative Music Loudness Annotations