Title :
A revisit to the class imbalance learning with linear support vector machine
Author :
Yang Fan ; Zheng Kai ; Li Qiang
Author_Institution :
Sch. of Inf. Sci. & Eng., Xiamen Univ., Xiamen, China
Abstract :
Existing re-sampling methods such as Synthetic minority over-sampling technique (SMOTE) and random under-sampling (RUS) perform unsatisfactorily in some imbalanced data, even outperformed by non-sampling method like standard linear support vector machine (SVM). In this paper, we employ support vectors to approximately estimate the ratio of two class instances close to the boundary, and then apply the ratio for re-sampling. Experimental results show that re-sampling using the boundary ratio will perform well on real imbalanced datasets and the standard linear SVM could have better performance than re-sampling methods. Therefore, in terms of data, balance or imbalance, should not be simply interpreted as the ratio of the overall number of two class instances, but should be interpreted as the ratio close to the boundary.
Keywords :
learning (artificial intelligence); sampling methods; support vector machines; RUS method; SMOTE method; SVM; boundary ratio; class imbalance learning; instance learning; linear support vector machine; random under-sampling method; resampling methods; synthetic minority over-sampling technique; Classification algorithms; Computers; Educational institutions; Glass; Vehicles; borderline ratio based sampling; imbalanced data; support vector machine;
Conference_Titel :
Computer Science & Education (ICCSE), 2014 9th International Conference on
Conference_Location :
Vancouver, BC
Print_ISBN :
978-1-4799-2949-8
DOI :
10.1109/ICCSE.2014.6926515