DocumentCode
3494862
Title
Uncertainty sampling methods for selecting datasets in active meta-learning
Author
Prudêncio, Ricardo B C ; Soares, Carlos ; Ludermir, Teresa B.
Author_Institution
Center of Inf., Fed. Univ. of Pernambuco, Recife, Brazil
fYear
2011
fDate
July 31 2011-Aug. 5 2011
Firstpage
1082
Lastpage
1089
Abstract
Several meta-learning approaches have been developed for the problem of algorithm selection. In this context, it is of central importance to collect a sufficient number of datasets to be used as meta-examples in order to provide reliable results. Recently, some proposals to generate datasets have addressed this issue with successful results. These proposals include datasetoids, which is a simple manipulation method to obtain new datasets from existing ones. However, the increase in the number of datasets raises another issue: in order to generate meta-examples for training, it is necessary to estimate the performance of the algorithms on the datasets. This typically requires running all candidate algorithms on all datasets, which is computationally very expensive. In a recent paper, active meta-learning has been used to address this problem. An uncertainty sampling method for the k-NN algorithm using a least confidence score based on a distance measure was employed. Here we extend that work, namely by investigating three hypotheses: 1) is there advantage in using a frequency-based least confidence score over the distance-based score? 2) given that the meta-learning problem used has three classes, is it better to use a margin-based score? and 3) given that datasetoids are expected to contain some noise, are better results achieved by starting the search with all datasets already labeled? Some of the results obtained are unexpected and should be further analyzed. However, they confirm that active meta-learning can significantly reduce the computational cost of meta-learning with potential gains in accuracy.
Keywords
data analysis; learning (artificial intelligence); sampling methods; active meta-learning; algorithm selection problem; dataset selection; datasetoids; distance measure; frequency-based least confidence score; k-NN algorithm; margin-based score; uncertainty sampling method; Computational efficiency; Entropy; Machine learning; Machine learning algorithms; Sampling methods; Training; Uncertainty;
fLanguage
English
Publisher
ieee
Conference_Titel
Neural Networks (IJCNN), The 2011 International Joint Conference on
Conference_Location
San Jose, CA
ISSN
2161-4393
Print_ISBN
978-1-4244-9635-8
Type
conf
DOI
10.1109/IJCNN.2011.6033343
Filename
6033343
Link To Document