Author/Authors
çelik, özer eskişehir osmangazi üniversitesi - fen-edebiyat fakültesi - matematik ve bilgisayar bilimleri bölümü, ESKİŞEHİR, turkey , kaplan, gürkan eskişehir osmangazi üniversitesi - fen-edebiyat fakültesi - matematik ve bilgisayar bilimleri bölümü, ESKİŞEHİR, turkey
Title Of Article
Text Classification Study on SMS Data Using Resampling Techniques
شماره ركورد
44934
Abstract
SMS is one of the important tools that mobile devices users use in their communication. Today, most of the information received by users is the source of mobile phones. With the advances in technology, the content of the messages coming to mobile phones is spread over a wide area and whether or not they come from the desired source is an important issue. The lack of Turkish studies in text classification studies is noteworthy. In this study, the messages received from a large number of users phones were examined and brought together through various improvement stages such as data preprocessing. After that, the existing message contents were examined by applying text classification by machine learning techniques. The data obtained are divided into 3 different categories as normal, advertising and spam. In order to stabilize the unbalanced data set, classification performances were examined by applying Synthetic Minority Oversampling Technique (SMOTE), Condensed Nearest Neighbor (CNN) Undersampling Technique and Random Undersampling Technique (RUS). As a result of the study performed on the data set containing 4203 SMS, the best classifications (according to OACC value) were Logistic Regression with 80.1% in SMOTE, XGBoost with 62.1% in CNN and Logistic Regression with 73.8% in RUS.
From Page
433
NaturalLanguageKeyword
Text Classification , Machine learning , Artificial intelligence , Smote , SMS
JournalTitle
Erciyes University Journal Of The Institute Of Science and Technology
To Page
442
JournalTitle
Erciyes University Journal Of The Institute Of Science and Technology
Link To Document