• DocumentCode
    117915
  • Title

    Topic model allocation of conversational dialogue records by Latent Dirichlet Allocation

  • Author

    Jui-Feng Yeh ; Chen-Hsien Lee ; Yi-Shiuan Tan ; Liang-Chih Yu

  • Author_Institution
    Dept. of Comput. & Inf. Sci., Nat. Chia-Yi Univ., Chiayi, Taiwan
  • fYear
    2014
  • fDate
    9-12 Dec. 2014
  • Firstpage
    1
  • Lastpage
    4
  • Abstract
    The topic information of conversational content is important for continuation with communication, so topic detection and tracking is one of important research. Due to there are many topic transform occurring frequently in long time communication, and the conversation maybe have many topics, so it´s important to detect different topics in conversational content. This paper detects topic information by using agglomerative clustering of utterances and Dynamic Latent Dirichlet Allocation topic model, uses proportion of verb and noun to analyze similarity between utterances and cluster all utterances in conversational content by agglomerative clustering algorithm. The topic structure of conversational content is friability, so we use speech act information and gets the hypernym information by E-HowNet that obtains robustness of word categories. Latent Dirichlet Allocation topic model is used to detect topic in file units, it just can detect only one topic if uses it in conversational content, because of there are many topics in conversational content frequently, and also uses speech act information and hypernym information to train the latent Dirichlet allocation models, then uses trained models to detect different topic information in conversational content. For evaluating the proposed method, support vector machine is developed for comparison. According to the experimental results, we can find the proposed method outperforms the approach based on support vector machine in topic detection and tracking in spoken dialogue.
  • Keywords
    document handling; natural language processing; pattern clustering; support vector machines; agglomerative utterance clustering; conversational dialogue records; dynamic latent Dirichlet allocation topic model; hypernym information; speech act information; support vector machine; topic detection; topic model allocation; topic tracking; Blogs; Computational modeling; Data models; Probability; Resource management; Semantics; Support vector machines; Conversational dialogue; latent Dirichlet allocation; spoken language processing; topic detection and tracking;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Asia-Pacific Signal and Information Processing Association, 2014 Annual Summit and Conference (APSIPA)
  • Conference_Location
    Siem Reap
  • Type

    conf

  • DOI
    10.1109/APSIPA.2014.7041546
  • Filename
    7041546