DocumentCode :
1037780
Title :
Semantic Annotation and Retrieval of Music and Sound Effects
Author :
Turnbull, Douglas ; Barrington, Luke ; Torres, David ; Lanckriet, Gert
Author_Institution :
Dept. of Comput. Sci. & Eng., Univ. of California at San Diego, La Jolla, CA
Volume :
16
Issue :
2
fYear :
2008
Firstpage :
467
Lastpage :
476
Abstract :
We present a computer audition system that can both annotate novel audio tracks with semantically meaningful words and retrieve relevant tracks from a database of unlabeled audio content given a text-based query. We consider the related tasks of content-based audio annotation and retrieval as one supervised multiclass, multilabel problem in which we model the joint probability of acoustic features and words. We collect a data set of 1700 human-generated annotations that describe 500 Western popular music tracks. For each word in a vocabulary, we use this data to train a Gaussian mixture model (GMM) over an audio feature space. We estimate the parameters of the model using the weighted mixture hierarchies expectation maximization algorithm. This algorithm is more scalable to large data sets and produces better density estimates than standard parameter estimation techniques. The quality of the music annotations produced by our system is comparable with the performance of humans on the same task. Our ldquoquery-by-textrdquo system can retrieve appropriate songs for a large number of musically relevant words. We also show that our audition system is general by learning a model that can annotate and retrieve sound effects.
Keywords :
Gaussian processes; audio databases; audio signal processing; content-based retrieval; music; Gaussian mixture model; audio annotation; audio feature space; computer audition system; content-based retrieval; music retrieval; parameter estimation; query-by-text system; semantic annotation; sound effects; text-based query; Audio databases; Content based retrieval; Humans; Information analysis; Instruments; Music information retrieval; Natural languages; Parameter estimation; Spatial databases; Vocabulary; Audio annotation and retrieval; music information retrieval; semantic music analysis;
fLanguage :
English
Journal_Title :
Audio, Speech, and Language Processing, IEEE Transactions on
Publisher :
ieee
ISSN :
1558-7916
Type :
jour
DOI :
10.1109/TASL.2007.913750
Filename :
4432652
Link To Document :
بازگشت