Main Article Content

Elsya Sabrina Asmita Simorangkir
Efori Bu'ulolo

Abstract

K-Means is one of the most widely used clustering algorithms because of its simplicity and computational efficiency. However, its performance often decreases when handling non-linear data due to the assumption that all attributes contribute equally to the distance calculation process. This study proposes a Variance-Weighted Distance Metrics K-Means (VWDM-KMeans) method that assigns attribute weights based on variance values to improve clustering quality. The proposed approach consists of Min-Max Normalization, variance calculation, weight generation, and integration of variance-based weights into the distance metric used by K-Means. Experiments were conducted on a non-linear dataset containing 103 records and 3 attributes (x, y, and z) with K = 3 clusters. The generated attribute weights were 0.3207, 0.3342, and 0.3451 for attributes x, y, and z, respectively. The performance of VWDM-KMeans was compared with conventional K-Means and K-Medoids using the number of iterations, Sum of Squared Errors (SSE), and Silhouette Score (SS). The results showed that VWDM-KMeans converged in 5 iterations, compared to 6 iterations for K-Means and 3 iterations for K-Medoids. In terms of cluster compactness, VWDM-KMeans achieved the lowest SSE value of 2.7932, outperforming K-Means (8.2429) and K-Medoids (8.9602). Furthermore, VWDM-KMeans obtained a Silhouette Score of 0.4854, equal to K-Means and higher than K-Medoids (0.4696). These findings demonstrate that incorporating variance-based attribute weighting into the distance calculation process improves cluster compactness while maintaining cluster separation quality and stability. Therefore, VWDM-KMeans can serve as an effective and computationally efficient alternative for clustering non-linear data.

Downloads

Download data is not yet available.

Article Details

How to Cite
Simorangkir, E. S. A., & Bu’ulolo, E. (2026). Improving K-Means clustering performance on non-linear data using variance-weighted distance metrics. Journal of Intelligent Decision Support System (IDSS), 9(2), 114-123. https://doi.org/10.35335/idss.v9i2.366
References
Bu’ulolo;Efori, Mesran, Hasibuan;Nelly Astuti, Utomo;Aripin;Soeb, Putro Utomo, S. (2023). Big Data Analysis dengan Phyton untuk Perguruan Tinggi (I).
Bu’ulolo, E., Sihombing, P., Sutarman, & Budiman, M. A. (2025). Variance-Weighted Centroid: A Centroid Estimation Approach for High-Dimensional Data Clustering. 2025 Tenth International Conference on Informatics and Computing (ICIC), 1–7. https://doi.org/10.1109/ICIC68054.2025.11309407
Bu’ulolo, E., Sihombing, P., Sutarman, S., & Budiman, M. (2026). K-Cube Consensus Clustering with Centroid Improvement and Variance-Based Metrics on High-Dimensional Data. Journal of Applied Data Sciences, 7(2), 1440–1454. https://doi.org/10.47738/jads.v7i2.1209
Du, X. (2023). A Robust and High-Dimensional Clustering Algorithm Based on Feature Weight and Entropy. Entropy, 25(3). https://doi.org/10.3390/e25030510
et al., E. B. (2022). A Review of Clustering Algorithms: Comparison of DBSCAN and K-Mean with Oversampling and t-SNE. Recent Patents on Engineering, 16(2). https://doi.org/10.2174/1872212115666210208222231
et al., L. S. (2023). Reproducible Clustering with Non-Euclidean Distances: A Simulation and Case Study. International Journal of Data Science and Analytics.
Galis, F., & Onchis, D. (2025). Refining Filter Global Feature Weighting for Fully-Unsupervised Clustering.
Golzari Oskouei, A., Balafar, M. A., & Motamed, C. (2021). FKMAWCW: Categorical fuzzy k-modes clustering with automated attribute-weight and cluster-weight learning. Chaos, Solitons & Fractals, 153, 111494. https://doi.org/https://doi.org/10.1016/j.chaos.2021.111494
Golzari Oskouei, A., Hashemzadeh, M., Asheghi, B., & Balafar, M. A. (2021). CGFFCM: Cluster-weight and Group-local Feature-weight learning in Fuzzy C-Means clustering algorithm for color image segmentation. Applied Soft Computing, 113, 108005. https://doi.org/https://doi.org/10.1016/j.asoc.2021.108005
H. Irwandi O. S. Sitompul, & Sutarman. (2023). K-Means Performance Optimization Using Rank Order Centroid (ROC) and Braycurtis Distance. Sinkron, 7(2). https://doi.org/10.33395/sinkron.v7i2.11371
Ikotun, A. M., Almutari, M. S., & Ezugwu, A. E. (2021). K-Means-Based Nature-Inspired Metaheuristic Algorithms for Automatic Data Clustering Problems: Recent Advances and Future Directions. Applied Sciences, 11(23). https://doi.org/10.3390/app112311246
Ikotun, A. M., Ezugwu, A. E., Abualigah, L., Abuhaija, B., & Heming, J. (2023). K-means clustering algorithms: A comprehensive review, variants analysis, and advances in the era of big data. Information Sciences, 622, 178–210. https://doi.org/10.1016/j.ins.2022.11.139
Irwandi, H., Sitompul, O. S., & Sutarman, S. (2022). K-Means Performance Optimization Using Rank Order Centroid (ROC) And Braycurtis Distance. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 6(2), 472–478. https://doi.org/10.33395/sinkron.v7i2.11371
Khan, A. A., Bashir, M. S., Batool, A., Raza, M. S., & Bashir, M. A. (2024). K-Means Centroids Initialization Based on Differentiation Between Instances Attributes. International Journal of Intelligent Systems, 2024(1). https://doi.org/10.1155/2024/7086878
Peng, J., Wang, D., & Wang, S. (2021). Feature-weighted distance metric learning for clustering. Pattern Recognition, 114, 107867. https://doi.org/10.1016/j.patcog.2021.107867
Pramudya, R. I., Kurniawan, T. A., Candra, M. H., Onn, C. W., & Dissanayake, K. K. (2026). Advancing unsupervised clustering: A systematic review of hybrid K-means and metaheuristic optimization algorithms in data mining. Computer Science Review, 61, 100972. https://doi.org/https://doi.org/10.1016/j.cosrev.2026.100972
Romanuke, V. V. (2023). Random Centroid Initialization for Improving Centroid-Based Clustering. Decision Making: Applications in Management and Engineering, 6(2), 734–746. https://doi.org/10.31181/dmame622023742
Simorangkir, E. S. A., Siahaan, A. P. U., Marlina, L., Nasution, D., & Sitorus, Z. (2024). Deteksi Outlier Hasil Clustering Algoritma K-Medoids Menggunakan Metode Boxplot Pada Data KIP Kuliah. Journal of Computer System and Informatics (JoSYC), 5(4), 893–902. https://doi.org/10.47065/josyc.v5i4.5479
Sinaga, K. P., Hussain, I., & Yang, M. S. (2021). Entropy K-Means Clustering with Feature Reduction under Unknown Number of Clusters. IEEE Access, 9, 67736–67751. https://doi.org/10.1109/ACCESS.2021.3077622
Syahputra, N., Zarlis, M., & Efendi, S. (2022). Seleksi Fitur Menggunakan Eigen Vector Untuk Peningkatan Kinerja K-Means Clustering Dalam Pengelompokan Data. Building of Informatics, Technology and Science (BITS), 4(2 SE-Articles). https://doi.org/10.47065/bits.v4i2.2022
Thrun, M. C. (2021). Distance-based clustering challenges for unbiased benchmarking studies. Scientific Reports, 11(1), 18988. https://doi.org/10.1038/s41598-021-98126-1
Wu, Z., Wang, B., & Li, C. (2022). A new robust fuzzy clustering framework considering different data weights in different clusters. Expert Systems with Applications, 206, 117728. https://doi.org/https://doi.org/10.1016/j.eswa.2022.117728
Zhang, R.-L., & Liu, X.-H. (2023). A Novel Hybrid High-Dimensional PSO Clustering Algorithm Based on the Cloud Model and Entropy. Applied Sciences, 13(3), 1246.