DocumentCode
2694757
Title
Web image selection with PLSA
Author
Yanai, Keiji
Author_Institution
Dept. of Comput. Sci., Univ. of Electro-Commun., Chofu
fYear
2008
fDate
June 23 2008-April 26 2008
Firstpage
1373
Lastpage
1376
Abstract
In this paper, we propose a new method to select relevant images to the given keywords from the images gathered from the Web. Our novel method is based on the probabilistic latent semantic analysis (PLSA) model, which is a generative probabilistic topic model. Firstly, we gather images related to the given keywords from the Web with Web search engines. Secondly, we choose pseudo-training images from them by simple heuristic HTML analysis, and train our PLSA-based probabilistic model with them. Finally, we select relevant images from all the gathered images with the learned model. The experimental results shows that the results by the proposed method is almost equivalent to the results by existing methods, although our method does not need to prepare negative training samples in advance unlike existing methods.
Keywords
Internet; hypermedia markup languages; image retrieval; search engines; HTML; World Wide Web; image selection; probabilistic latent semantic analysis; search engines; Computer science; HTML; Image analysis; Image databases; Image recognition; Large-scale systems; Object recognition; Search engines; Support vector machines; Training data; PLSA; Web image; bag-of-visual-words;
fLanguage
English
Publisher
ieee
Conference_Titel
Multimedia and Expo, 2008 IEEE International Conference on
Conference_Location
Hannover
Print_ISBN
978-1-4244-2570-9
Electronic_ISBN
978-1-4244-2571-6
Type
conf
DOI
10.1109/ICME.2008.4607699
Filename
4607699
Link To Document