Title :
Having a Blast: Meta-Learning and Heterogeneous Ensembles for Data Streams
Author :
Jan N. van Rijn;Geoffrey Holmes;Bernhard Pfahringer;Joaquin Vanschoren
Author_Institution :
Leiden Univ., Leiden, Netherlands
Abstract :
Ensembles of classifiers are among the best performing classifiers available in many data mining applications. However, most ensembles developed specifically for the dynamic data stream setting rely on only one type of base-level classifier, most often Hoeffding Trees. In this paper, we study the use of heterogeneous ensembles, comprised of fundamentally different model types. Heterogeneous ensembles have proven successful in the classical batch data setting, however they do not easily transfer to the data stream setting. We therefore introduce the Online Performance Estimation framework, which can be used in data stream ensembles to weight the votes of (heterogeneous) ensemble members differently across the stream. Experiments over a wide range of data streams show performance that is competitive with state of the art ensemble techniques, including Online Bagging and Leveraging Bagging. All experimental results from this work are easily reproducible and publicly available on OpenML for further analysis.
Keywords :
"Training","Estimation","Data models","Data mining","Bagging","Stacking","Predictive models"
Conference_Titel :
Data Mining (ICDM), 2015 IEEE International Conference on
DOI :
10.1109/ICDM.2015.55