• DocumentCode
    178653
  • Title

    Leveraging repetition for improved automatic lyric transcription in popular music

  • Author

    McVicar, Matt ; Ellis, Daniel P. W. ; Goto, Misako

  • Author_Institution
    Nat. Inst. of Adv. Ind. Sci. & Technol., Tsukuba, Japan
  • fYear
    2014
  • fDate
    4-9 May 2014
  • Firstpage
    3117
  • Lastpage
    3121
  • Abstract
    Transcribing lyrics from musical audio is a challenging research problem which has not benefited from many advances made in the related field of automatic speech recognition, owing to the prevalent musical accompaniment and differences between the spoken and sung voice. However, one aspect of this problem which has yet to be exploited by researchers is that significant portions of the lyrics will be repeated throughout the song. In this paper we investigate how this information can be leveraged to form a consensus transcription with improved consistency and accuracy. Our results show that improvements can be gained using a variety of techniques, and that relative gains are largest under the most challenging and realistic experimental conditions.
  • Keywords
    music; speech recognition; automatic speech recognition; consensus transcription; improved automatic lyric transcription; musical accompaniment; musical audio; popular music; repetition leveraging; spoken voice; sung voice; Accuracy; Hidden Markov models; Mel frequency cepstral coefficient; Speech; Speech processing; Speech recognition; Automatic Lyric Recognition; Automatic Speech Recognition; Music Information Retrieval;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on
  • Conference_Location
    Florence
  • Type

    conf

  • DOI
    10.1109/ICASSP.2014.6854174
  • Filename
    6854174