• DocumentCode
    3494862
  • Title

    Uncertainty sampling methods for selecting datasets in active meta-learning

  • Author

    Prudêncio, Ricardo B C ; Soares, Carlos ; Ludermir, Teresa B.

  • Author_Institution
    Center of Inf., Fed. Univ. of Pernambuco, Recife, Brazil
  • fYear
    2011
  • fDate
    July 31 2011-Aug. 5 2011
  • Firstpage
    1082
  • Lastpage
    1089
  • Abstract
    Several meta-learning approaches have been developed for the problem of algorithm selection. In this context, it is of central importance to collect a sufficient number of datasets to be used as meta-examples in order to provide reliable results. Recently, some proposals to generate datasets have addressed this issue with successful results. These proposals include datasetoids, which is a simple manipulation method to obtain new datasets from existing ones. However, the increase in the number of datasets raises another issue: in order to generate meta-examples for training, it is necessary to estimate the performance of the algorithms on the datasets. This typically requires running all candidate algorithms on all datasets, which is computationally very expensive. In a recent paper, active meta-learning has been used to address this problem. An uncertainty sampling method for the k-NN algorithm using a least confidence score based on a distance measure was employed. Here we extend that work, namely by investigating three hypotheses: 1) is there advantage in using a frequency-based least confidence score over the distance-based score? 2) given that the meta-learning problem used has three classes, is it better to use a margin-based score? and 3) given that datasetoids are expected to contain some noise, are better results achieved by starting the search with all datasets already labeled? Some of the results obtained are unexpected and should be further analyzed. However, they confirm that active meta-learning can significantly reduce the computational cost of meta-learning with potential gains in accuracy.
  • Keywords
    data analysis; learning (artificial intelligence); sampling methods; active meta-learning; algorithm selection problem; dataset selection; datasetoids; distance measure; frequency-based least confidence score; k-NN algorithm; margin-based score; uncertainty sampling method; Computational efficiency; Entropy; Machine learning; Machine learning algorithms; Sampling methods; Training; Uncertainty;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Neural Networks (IJCNN), The 2011 International Joint Conference on
  • Conference_Location
    San Jose, CA
  • ISSN
    2161-4393
  • Print_ISBN
    978-1-4244-9635-8
  • Type

    conf

  • DOI
    10.1109/IJCNN.2011.6033343
  • Filename
    6033343