Clustering Pospay User Reviews Using Rule-Based Hybrid K-Means and TF-IDF
DOI:
10.33395/sinkron.v10i4.16556Keywords:
Text mining, Rule-Based Hybrid K-Means, TF-IDF Bigram, User Reviews, PospayAbstract
Pospay is a digital financial services application created by PT Pos Indonesia, which offers a variety of financial transaction services. The huge volume of unstructured data and the exponential growth of the number of users’ assessments in Google Play Store makes it impossible to be analyzed manually. This research aims to classify Pospay customer evaluation to identify the main problems experienced by the user and provide recommendation in upgrading application services. This work adopts the Knowledge Discovery in Databases (KDD) technique. The inputs include 11000 user evaluations, of which 10943 are maintained after pre-processing. The input text has been vectorized using TF-IDF Bigram and the ideal number of clusters has been calculated using the Elbow Method and Silhouette Score. Then, a Rule-Based Hybrid K-Means technique was used, which integrated K-Means++ clustering with rule-based refinement to enhance the interpretability of the clusters. The findings produced five primary clusters that are related to verification and identity, system and error, login and account, transaction and balance, and positive reviews. Authentication and identification and transaction and balance were the most talked-about issues among users from these countries, accounting for 31.3% and 26.8% of discussions respectively. PCA visualization and word cloud analysis further supported the interpretation of each cluster. Overall, the proposed approach effectively grouped user reviews into meaningful topics and can assist developers in identifying service priorities to improve the quality and reliability of the Pospay application.
Downloads
References
Annas, M., & Wahab, S. N. (2023). Data Mining Methods: K-Means Clustering Algorithms. International Journal of Cyber and IT Service Management (IJCITSM), 3(1), 40–47. https://iiast.iaic-publisher.org/ijcitsm/index.php/IJCITSM/article/view/122
Dąbrowski, J., Letier, E., Perini, A., & Susi, A. (2022). Analysing app reviews for software engineering: a systematic literature review. Empirical Software Engineering, 27(2). https://doi.org/10.1007/s10664-021-10065-7
Handayani, F. D., & Rosyida, I. (2023). Clustering Review Pengguna Aplikasi Zenius pada Layanan Google Play Store Menggunakan Metode DBSCAN dan HDBSCAN. Emerging Statistics and Data Science Journal, 1(2), 178–191.
June, V. N., Widyawati, F., Dawod, A. Y., & Santoso, H. A. (2025). K-Means Clustering Optimization of Toddler Malnutrition Status Using Elbow Method. Journal of Informatics and Web Engineering EISSN:, 4(2).
Khairuna, R., Nurdin, & Ar Razi. (2025). Analisis Perbandingan Algoritma K-Means Dan K-Medoids Untuk Klusterisasi Teks Ulasan Pada Aplikasi Cookpad. Rabit : Jurnal Teknologi Dan Sistem Informasi Univrab, 10(2), 445–458. https://doi.org/10.36341/rabit.v10i2.6131
Kim, G.-Y., & Han, S. (2022). User Review Analysis of English Learning Applications on Google Play Store Using Text-Mining. Journal of Digital Contents Society, 23(10), 1901–1908. https://doi.org/10.9728/dcs.2022.23.10.1901
Mahendra, K., Rahman, Z. S., Supendar, H., & Fahlapi, R. (2026). Pemetaan Tema Keluhan Pengguna Aplikasi Cek Bansos Menggunakan K- Means dan TF-IDF Berbasis Ulasan Google Play Store. JOURNAL KOMPUTER TERKNOLOGI INFORMASI SISTEM KOMPUTER(JUKTISI9), 5(1), 550–559.
Memon, Z. A., Munawar, N., & Kamal, M. (2023). App store mining for feature extraction: analyzing user reviews. Acta Scientiarum - Technology, 46, 1–16. https://doi.org/10.4025/actascitechnol.v46i1.62867
Mohammed, A. A., Sumari, P., & Attabi, K. (2024). Hybrid K-means and Principal Component Analysis (PCA) for Diabetes Prediction: International Journal of Computing and Digital Systems, 15(1), 1719–1728. https://doi.org/10.12785/ijcds/1501121
Mukti, B. P., Hariguna, T., & Tahyudin, I. (2025). Model Klastering Hybrid Menggunakan Inisialisasi K-means++ dan Algoritma Optimasi Grey Wolf. Jurnal Sistem Dan Teknologi Informasi, 13(2), 286–298. https://doi.org/10.26418/justin.v13i2.88211
Nasser, F. K., & Behadili, S. F. (2022). A Review of Data Mining and Knowledge Discovery Approaches for Bioinformatics ةيجهلهيابلا تامهمعملا لاجم يف قئاقحلا فاشتكاو تانايبلا نيدعت ةقيرطل ضا رعتسا. Iraqi Journal of Science, 63(7), 3169–3188. https://doi.org/10.24996/ijs.2022.63.7.37
Nunkaew, W., Ahmed, M., & Nakfon, P. (2026). A Machine Learning-Based Hybrid K-means and Fuzzy Inference System with Rule-Based Fine-Tuning for Sustainable Supplier Segmentation. In AICCC 2025 - 2025 8th Artificial Intelligence and Cloud Computing Conference (Vol. 1, Number 1). Association for Computing Machinery. https://doi.org/10.1145/3789982.3789986
Nurul, K., Djati, I., & Faiza, N. (2023). Identifying Improvement Strategic from User Application Reviews Group Using K-Means Clustering and TF-IDF Weighting. International Journal of Artificial Intelegence Research, 7(2), 152–159.
Onumanyi, A. J., Molokomme, D. N., Isaac, S. J., & Abu-mahfouz, A. M. (2022). AutoElbow : An Automatic Elbow Detection Method for Estimating the Number of Clusters in a Dataset. Applied Sciences, 12(15), 7515. https://doi.org/https://doi.org/10.3390/app12157515
Pamput, J. P., Muthmainnah, A. R., Risal, A. A. N., & Surianto, D. F. (2025). K-Means++ and TF-IDF for Grouping Library Books by Topic. Paradigma - Jurnal Komputer Dan Informatika, 27(2), 74–82. https://doi.org/10.31294/p.v27i2.8272
Pamungkas, M. D., & Februariyanti, H. (2022). Penerapan Algoritma K-Means Clustering Untuk Mengelompokan Data Review Barang Pada E-Commerce Lazada. SemanTIK, 8(2), 99. https://doi.org/10.55679/semantik.v8i2.29058
Prasetyadi, A., Nugroho, B., & Tohari, A. (2022). A Hybrid K-Means Hierarchical Algorithm for Natural Disaster Mitigation Clustering. Journal of Information and Communication Technology, 2(2), 175–200. https://doi.org/https://doi.org/10.32890/jict2022.21.2.2
Ramadhini, N., Annisaa, S. A., Putri, A. S., Tania, K. D., & Rifai, A. (2026). KLASTERISASI PORTAL BERITA ONLINE MENGGUNAKAN ALGORITMA K-MEANS DENGAN TEKNIK KNOWLEDGE DISCOVERY IN DATABASE. JATI (Jurnal Mahasiswa Teknik Informatika), 10(3), 4079–4087. https://doi.org/https://doi.org/10.36040/jati.v10i3.18100
Sudrajat, R., Hadiana, A. I., & Melina. (2025). Evaluasi Kualitas Klaster Wilayah Rawan Bencana Menggunakan K- Means dengan Silhouette dan Elbow Method. Jurnal Algoritma, 22(2), 127–139. https://doi.org/10.33364/algoritma/v.22-2.2379
Ulummuddin, I., Sari, A. P., & Swari, M. H. P. (2024). PEMANFAATAN DATA ULASAN PENGGUNA UNTUK MEMBANGUN SISTEM KLASTERISASI BERDASARKAN PAIN POINTS MENGGUNAKAN ALGORITMA K-MEANS. Jurnal Teknologi Terpadu, 10(1), 70–76. https://doi.org/https://doi.org/10.54914/jtt.v10i1.1252
Vania, P., & Sari, B. N. (2023). Perbandingan Metode Elbow dan Silhouette untuk Penentuan Jumlah Klaster yang Optimal pada Clustering Produksi Padi menggunakan Algoritma K-Means. Jurnal Ilmiah Wahana Pendidikan, 9(21), 547–558. https://doi.org/https://doi.org/10.5281/zenodo.10081332
Zhu, J., Huang, S., Shi, Y., Wu, K., & Wang, Y. (2022). A Method of K-Means Clustering Based on TF-IDF for Software Requirements Documents Written in Chinese Language. IEICE Transactions on Information and Systems, 105(4), 736–754. https://doi.org/10.1587/transinf.2021EDP7144
Downloads
How to Cite
Issue
Section
License
Copyright (c) 2026 Suci Tamaro Siahaan, Christian Dwi Suhendra, Lion Ferdinand Marini

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.






















Moraref
PKP Index
Indonesia OneSearch
OCLC Worldcat
Index Copernicus
Scilit
