Random Forest Implementation for Indonesian Coffee Identification Based on Digital Image Feature Extraction

Authors

  • Widya Lelisa Army Doctoral Program in Informatics, Universitas Ahmad Dahlan, Yogyakarta, Indonesia 55191
  • Abdul Fadlil Doctoral Program in Informatics, Universitas Ahmad Dahlan, Yogyakarta, Indonesia 55191
  • Sunardi Doctoral Program in Informatics, Universitas Ahmad Dahlan, Yogyakarta, Indonesia 55191

DOI:

10.33395/sinkron.v10i4.16622

Keywords:

Keywords: coffee bean classification; ensemble learning; feature importance; GLCM; random forest

Abstract

Accurate identification of coffee types is a critical challenge in Indonesia's coffee industry, which relies heavily on subjective human visual inspection. This study implements the Random Forest ensemble learning algorithm to classify three commercially important coffee species  Arabica (Coffea arabica), Liberica (Coffea liberica), and Robusta (Coffea canephora) using digital image features. The dataset consists of 1,913 real coffee bean images from Roboflow Universe (Coffee Bean Type v1, CC BY 4.0), split into train (1,530), valid (287), and test (96). Feature extraction produces an 11-dimensional vector combining five Gray Level Co-occurrence Matrix (GLCM) texture features (Contrast, Energy, Homogeneity, Correlation, Dissimilarity) computed at four angles and six RGB color statistical features (mean and standard deviation per channel). The model is trained on 1,817 combined train-valid samples with optimal hyperparameters (n_estimators=200, max_depth=None, criterion=gini) selected via Grid Search with 5-fold stratified cross-validation. Evaluation on 96 held-out test samples yields accuracy 98.96%, macro precision 99.05%, macro recall 98.67%, macro F1-score 98.84%, and ROC-AUC 0.9996, with only one misclassification. Cross-validation confirms model stability at 98.07%±0.60%. Feature importance analysis identifies GLCM Energy (21.98%) and Homogeneity (14.57%) as the most discriminative features. The Robusta class achieves perfect classification (100% all metrics). Manual verification confirms that the reported metrics are consistent with the confusion matrix. These results demonstrate that Random Forest with GLCM and RGB color statistical features is effective for image-based coffee classification, achieving higher accuracy than several prior KNN and neural-network-based studies, although those studies differ in task, dataset, and class composition and are therefore not directly comparable.

GS Cited Analysis

Downloads

Download data is not yet available.

References

Breiman, L. (2001). Random Forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324

Direktorat Jenderal Perkebunan, Kementerian Pertanian RI. (2023). Statistik Perkebunan Jilid I 2022–2024. Jakarta: Ditjenbun Kementan RI.

Fatchurrachman, A., & Udjulawa, D. (2021). Identifikasi penyakit daun kopi menggunakan citra digital. Jurnal Ilmu Komputer dan Bisnis, 14(1), 66–73. https://doi.org/10.47927/jikbv14i1.622

Gonzalez, R. C., & Woods, R. E. (2018). Digital image processing (4th ed.). New York: Pearson Education.

Hakim, L., Janandi, C., & Cenggoro, T. W. (2020). Automated coffee roast degree classification using MobileNetV2. Proceedings of the International Conference on Informatics, Jakarta.

Haralick, R. M., Shanmugam, K., & Dinstein, I. (1973). Textural features for image classification. IEEE Transactions on Systems, Man, and Cybernetics, 3(6), 610–621. https://doi.org/10.1109/TSMC.1973.4309314

Iskandar, J. (2024). Coffee Bean Type Dataset v1. Roboflow Universe. Retrieved from https://universe.roboflow.com/jaya-iskandar/coffee-bean-type [CC BY 4.0]

Liaw, A., & Wiener, M. (2002). Classification and regression by randomForest. R News, 2(3), 18–22.

Metha, H. S., Kusrini, K., & Ariatmanto, D. (2024). Classification of types roasted coffee beans using convolutional neural network method. SinkrOn: Jurnal dan Penelitian Teknik Informatika, 8(2), 846–851. https://doi.org/10.33395/sinkron.v8i2.13517

Mehta, B., Bharany, S., Elkamchouchi, D. H., Khalaf, O. I., Garsamo, E. T., Hariharan, U., & Arora, N. (2025). EffResViT-SE FusionNet: A hybrid deep learning framework for accurate classification of coffee leaf diseases. Food Science and Nutrition. https://doi.org/10.1002/fsn3.71311

Motta, I. V. C., Vuillerme, N., Pham, H. H., & De Figueiredo, F. A. P. (2025). Machine learning techniques for coffee classification: A comprehensive review of scientific research. Artificial Intelligence Review, 58. https://doi.org/10.1007/s10462-024-11004-w

Murphy, K. P. (2012). Machine learning: A probabilistic perspective. Cambridge: MIT Press.

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., … Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.

Rahardjo, B. (2015). Panduan budidaya dan pengolahan kopi Arabika dan Robusta. Jakarta: Penebar Swadaya.

Soh, L. K., & Tsatsoulis, C. (1999). Texture analysis of SAR sea ice imagery using gray level co-occurrence matrices. IEEE Transactions on Geoscience and Remote Sensing, 37(2), 780–795. https://doi.org/10.1109/36.752194

Triyono, S., Laksono, R. A., & Tusi, A. (2023). Identifikasi jenis kopi menggunakan sensor e-nose dengan metode jaringan syaraf tiruan backpropagation. Jurnal Ilmiah Rekayasa Pertanian dan Biosistem, 11(1), 1–12.

Ardah, H., Alrahhal, M., Abd-Elhafiez, W. M., & Trabay, D. (2025). Robust coffee plant disease classification using deep learning and advanced feature engineering techniques. PeerJ Computer Science, e3386. https://doi.org/10.7717/peerj-cs.3386

Downloads


Crossmark Updates

How to Cite

Army, W. L., Fadlil, A. ., & Sunardi, S. (2026). Random Forest Implementation for Indonesian Coffee Identification Based on Digital Image Feature Extraction. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(4), 2405-2415. https://doi.org/10.33395/sinkron.v10i4.16622