• Title of article

    Test Data Likelihood for PLSA Models

  • Author/Authors

    Thorsten، Brants نويسنده ,

  • Issue Information
    روزنامه با شماره پیاپی سال 2005
  • Pages
    -180
  • From page
    181
  • To page
    0
  • Abstract
    Probabilistic Latent Semantic Analysis (PLSA) is a statistical latent class model that has recently received considerable attention. In its usual formulation it cannot assign likelihoods to unseen documents. Furthermore, it assigns a probability of zero to unseen documents during training. We point out that one of the two existing alternative formulations of the Expectation-Maximization algorithms for PLSA does not require this assumption. However, even that formulation does not allow calculation ofthe actual likelihood values. We therefore derive a new test-data likelihood substitute for PLSA and compare it to three existing likelihood substitutes. An empirical evaluation shows that our new likelihood substitute produces the best predictions about accuracies in two different IR tasks and is therefore best suited to determine the number of EM steps when training PLSA models. The new likelihood measure and its evaluation also suggest that PLSA is not very sensitive to overfitting for the two tasks considered. This renders additions like tempered EM that especially address overfitting unnecessary.
  • Keywords
    paraboea rufescens , reprodutive biology , mirror image flowers , xishuangbanna , buzz pollination , enantiostyly , Gesneriaceae
  • Journal title
    INFORMATION RETRIEVAL
  • Serial Year
    2005
  • Journal title
    INFORMATION RETRIEVAL
  • Record number

    89781