Genetic Algorithm Ensemble Filter Methods on Kidney Disease Classification

Authors

  • Sharin Hazlin Huspi School of Computing, Faculty of Engineering, Universiti Teknologi Malaysia (UTM), 81310 Johor Bahru, Johor
  • Chong Ke Ting School of Computing, Faculty of Engineering, Universiti Teknologi Malaysia (UTM), 81310 Johor Bahru, Johor

DOI:

https://doi.org/10.11113/ijic.v11n2.345

Keywords:

Feature selection, genetic algorithm, random forest, k-nearest neighbor, XGBoost, support vector machine, naïve bayes

Abstract

Kidney failure will give effect to the human body, and it can lead to a series of seriously illness and even causing death. Machine learning plays important role in disease classification with high accuracy and shorter processing time as compared to clinical lab test. There are 24 attributes in the Chronic K idney Disease (CKD) clinical dataset, which is considered as too much of attributes. To improve the performance of the classification, filter feature selection methods used to reduce the dimensions of the feature and then the ensemble algorithm is used to identify the union features that selected from each filter feature selection. The filter feature selection that implemented in this research are Information Gain (IG), Chi-Squares, ReliefF and Fisher Score. Genetic Algorithm (GA) is used to select the best subset from the ensemble result of the filter feature selection. In this research, Random Forest (RF), XGBoost, Support Vector Machine (SVM), K-Nearest Neighbor (KNN) and Naïve Bayes classification techniques were used to diagnose the CKD. The features subset that selected are different and specialised for each classifier. By implementing the proposed method irrelevant features through filter feature selection able to reduce the burden and computational cost for the genetic algorithm. Then, the genetic algorithm able to perform better and select the best subset that able to improve the performance of the classifier with less attributes. The proposed genetic algorithm union filter feature selections improve the performance of the classification algorithm. The accuracy of RF, XGBoost, KNN and SVM can achieve to 100% and NB can achieve to 99.17%. The proposed method successfully improves the performance of the classifier by using less features as compared to other previous work.

Downloads

Published

2021-10-31

How to Cite

Huspi, S. H. ., & Ke Ting, C. . (2021). Genetic Algorithm Ensemble Filter Methods on Kidney Disease Classification. International Journal of Innovative Computing, 11(2), 73–80. https://doi.org/10.11113/ijic.v11n2.345

Issue

Section

Computer Science