Skripsi
KLASIFIKASI PENYAKIT HEPATITIS MENGGUNAKAN TEACHING-LEARNING BASED OPTIMIZATION DAN RANDOM FOREST
Hepatitis is one of the global health problems with a persistently high incidence rate and the potential to cause serious complications if not detected early. Meanwhile, the use of laboratory data for machine learning-based diagnosis still faces challenges such as missing values, imbalanced class distribution, and the limited application of optimal combinations of feature selection methods and classification algorithms on the Hepatitis C dataset from Kaggle. This study proposes a hepatitis disease classification model using the Random Forest algorithm optimized with the Teaching-Learning Based Optimization (TLBO) feature selection method. The dataset consists of 615 patient records with five diagnostic categories: Hepatitis A, Hepatitis B, Hepatitis C, Blood Donor, and Suspect Blood Donor. The preprocessing stages include missing value imputation using the median, outlier handling using the IQR-based capping method, Min-Max normalization, and data balancing using the Synthetic Minority Over-sampling Technique (SMOTE). The experimental results show that the combination of SMOTE, TLBO, and Random Forest with the number of learners (N) = 30 achieved the best performance, with an accuracy of 96.00% on the validation data and 95.16% on the test data, as well as a precision of 87.47%, recall of 76.00%, F1- score of 78.29%, and specificity of 95.56%. These findings demonstrate that TLBO is effective in improving the classification performance of Random Forest without reducing the model’s generalization ability and has the potential to be applied as an accurate and efficient decision support system for hepatitis diagnosis. Keywords: Classification, Hepatitis, Machine Learning, Teaching-Learning Based Optimization (TLBO), Random Forest, Feature Selection, Synthetic Minority OverSampling Technique (SMOTE)
No other version available