Comparison of Classification Methods for Email Spam Detection

Authors

  • Rifqi Arya Saputra Universitas Dian Nuswantoro
  • Fikri Budiman

DOI:

10.33395/sinkron.v10i4.16789

Keywords:

Email spam, Machine learning, Naïve Bayes, Random Forest, Spam detection, Support Vector Machine

Abstract

Email spam remains a persistent cybersecurity problem, since unsolicited messages waste user time and often deliver phishing or malicious content. Prior comparative studies of conventional spam classifiers rarely state clearly whether their preprocessing avoids leakage between training and test data, or whether reported metrics come from a held-out test set or from cross-validation. This study addresses that gap by comparing five conventional classifiers — Naïve Bayes, Logistic Regression, Decision Tree, Random Forest, and Support Vector Machine (SVM) — under a pipeline in which the 80:20 train-test split is performed first, TF-IDF is fit only on the training partition, and SMOTE oversampling is applied only to the training data, so the test set stays unseen during model development. A public Kaggle email dataset of 5,157 messages (4,516 legitimate, 641 spam) was used, with spam explicitly defined as the positive class. On the untouched test set, SVM achieved the best overall performance (98.55% accuracy, 98.40% precision, 90.44% recall, 94.25% F1-score), followed by Random Forest (98.16% accuracy, 97.56% precision), while Logistic Regression obtained the highest recall (95.59%). Comparing these results with five-fold cross-validation on the already-oversampled training data revealed a large optimistic gap — up to about 30 points in precision for Naïve Bayes — demonstrating why resampling must be repeated inside each fold rather than applied once beforehand. The findings show that reported performance depends strongly on how resampling interacts with data splitting, and provide a transparent, leakage-aware baseline for future spam-detection research.

GS Cited Analysis

Downloads

Download data is not yet available.

References

AbdulNabi, I., & Yaseen, Q. (2021). Spam Email Detection Using Deep Learning Techniques. Procedia Computer Science, 184, 853–858. https://doi.org/10.1016/j.procs.2021.03.107

Adnan, M., Imam, M. O., Javed, M. F., & Murtza, I. (2024). Improving spam email classification accuracy using ensemble techniques: A stacking approach. International Journal of Information Security, 23(1), 505–517. https://doi.org/10.1007/s10207-023-00756-1

Ahmed, N., Amin, R., Aldabbas, H., Koundal, D., Alouffi, B., & Shah, T. (2022). Machine Learning Techniques for Spam Detection in Email and IoT Platforms: Analysis and Research Challenges. Security and Communication Networks, 2022, 1–19. https://doi.org/10.1155/2022/1862888

Budiman, D., Zayyan, Z., Mardiana, A., & Mahrani, A. A. (2024). Email spam detection: A comparison of svm and naive bayes using bayesian optimization and grid search parameters. Journal of Student Research Exploration, 2(1), 53–64. https://doi.org/10.52465/josre.v2i1.260

Dhar, A., Anusha, K. O. V., Kataria, A., & Khan, M. A. (2023). Comparative Analysis of Deep Learning, SVM, Random Forest, and XGBoost for Email Spam Detection: A Socio- Network Analysis Approach. 2023 International Conference on Computing, Communication, and Intelligent Systems (ICCCIS), 701–707. https://doi.org/10.1109/ICCCIS60361.2023.10425771

Guo, Y., Mustafaoglu, Z., & Koundal, D. (2022). Spam Detection Using Bidirectional Transformers and Machine Learning Classifier Algorithms. Journal of Computational and Cognitive Engineering, 2(1), 5–9. https://doi.org/10.47852/bonviewJCCE2202192

Jáñez-Martino, F., Alaiz-Rodríguez, R., González-Castro, V., Fidalgo, E., & Alegre, E. (2023). A review of spam email detection: Analysis of spammer strategies and the dataset shift problem. Artificial Intelligence Review, 56(2), 1145–1173. https://doi.org/10.1007/s10462-022-10195-4

Kontsewaya, Y., Antonov, E., & Artamonov, A. (2021). Evaluating the Effectiveness of Machine Learning Methods for Spam Detection. Procedia Computer Science, 190, 479–486. https://doi.org/10.1016/j.procs.2021.06.056

Li, S., Li, Y., & Xu, H. (2024). Spam classification based on parallel optimized BERT. Applied and Computational Engineering, 41(1), 153–159. https://doi.org/10.54254/2755-2721/41/20230736

Nasreen, G., Murad Khan, M., Younus, M., Zafar, B., & Kashif Hanif, M. (2024). Email spam detection by deep learning models using novel feature selection technique and BERT. Egyptian Informatics Journal, 26, 100473. https://doi.org/10.1016/j.eij.2024.100473

Saeed, M. M., & Aghbari, Z. A. (2023). Survey on Deep Learning Approaches for Detection of Email Security Threat. Computers, Materials & Continua, 77(1), 325–348. https://doi.org/10.32604/cmc.2023.036894

Tian, Y., Dai, X., Li, Z., Guo, H., & Mao, X. (2025). Improving the accuracy of cybersecurity spam email detection using ensemble techniques: A stacking approach Machine learning for spam email detection. PLOS One, 20(9), e0331574. https://doi.org/10.1371/journal.pone.0331574

Wang, X. (2024). Spam Filtering in the Modern Era: A Review of Machine Learning, Deep Learning, and System Comparisons: Proceedings of the 2nd International Conference on Data Analysis and Machine Learning, 451–458. https://doi.org/10.5220/0013526000004619

Zavrak, S., & Yilmaz, S. (2023). Email spam detection using hierarchical attention hybrid deep learning method. Expert Systems with Applications, 233, 120977. https://doi.org/10.1016/j.eswa.2023.120977

Zhang, W., Yoshida, T., & Tang, X. (2011). A comparative study of TF*IDF, LSI and multi-words for text classification. Expert Systems with Applications, 38(3), 2758–2765. https://doi.org/10.1016/j.eswa.2010.08.066

Downloads


Crossmark Updates

How to Cite

Saputra, R. A., & Budiman, F. . (2026). Comparison of Classification Methods for Email Spam Detection. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(4), 2198-2206. https://doi.org/10.33395/sinkron.v10i4.16789