K-Means clustering is a method to grouping data based on the similarity of features and detect the hidden patterns in dataset. The dataset is from GARUDA Repository which contains raw data of PDF files. GARUDA dataset extraction process used static analysis method. The data extraction process produced twenty�one features using PDFiD. GARUDA dataset has a multi-class and imbalanced data, there…