• DocumentCode
    3704164
  • Title

    An Algorithm for Identifying the Learning Patterns in Big Data

  • Author

    Majed Farrash;Wenjia Wang

  • Volume
    2
  • fYear
    2015
  • Firstpage
    48
  • Lastpage
    55
  • Abstract
    Divide-and-Conquer is probably the most commonly used strategy to deal with a big data that is too big to be loaded into any computing system´s memory as a whole for analysis. It partitions such a big dataset into many smaller subsets that can be loaded into computer memory separately to induce models, which can be combined by machine learning ensemble methods. However, it is not clear that how the size of subsets may affect the learning performance of individual models and their ensemble. This paper proposes an ensemble based algorithm to quickly detect their relational patterns in terms of ensemble accuracy and the size of partitioned data subset. An ensemble framework of the algorithm is implemented and tested on 12 relatively big benchmark datasets. The experimental results indicate that it is able to identify the relation patterns accurately and efficiently in less than 10 steps. The identified patterns show that in most cases it is not necessary to use the whole big dataset for analysis as few smaller subsets are already sufficiently representative of the underlying problem, which is obviously a useful knowledge in big data analysis.
  • Keywords
    "Partitioning algorithms","Big data","Memory management","Prediction algorithms","Training","Bagging","Data mining"
  • Publisher
    ieee
  • Conference_Titel
    Trustcom/BigDataSE/ISPA, 2015 IEEE
  • Type

    conf

  • DOI
    10.1109/Trustcom.2015.561
  • Filename
    7345474