Cost-Aware Selective Text Classification with Calibrated Lightweight Models

Authors

  • Saif Alhusseny Department of Physics, College of Education for Pure Science, University of Mosul
  • Omar Alniemi Department of Computer Science, College of Education for Pure Science, University of Mosul
  • Alaa Saadoon Department of Computer Science, College of Education for Pure Science, University of Mosul
  • Hanaa Mahmood Department of Computer Science, College of Education for Pure Science, University of Mosul

DOI:

10.33395/sinkron.v10i4.16650

Keywords:

Confidence Threshold, Logistic Regression, Probability Calibration, Selective Prediction, Text Classification

Abstract

Reliable text classification requires calibrated confidence, abstention for uncertain cases, and explicit computational-cost reporting. This study evaluates cost-aware selective prediction for calibrated lightweight text classifiers on commodity CPU runtimes. The goal is to find a light-weight classifier that is both deployment friendly and achieves a balance between predictive accuracy, calibration, selective reliability, and computational efficiency. The pipeline trains Multinomial Naive Bayes, Logistic Regression, and LightGBM on TF-IDF features, applies post-hoc probability calibration, and accepts only predictions above a validation-selected confidence threshold. Calibrated Logistic Regression achieves the best practical accuracy-calibration-selective-risk-latency-review-budget trade-off in SMS spam, IMDb sentiment and AGNews topic classification experiments. Calibrated Logistic Regression achieves 0.9857 accuracy and 0.0070 ECE on SMS, 0.8718 accuracy and 0.0162 ECE on IMDb, and 0.8998 accuracy on AGNews. On IMDb, Logistic Regression sacrifices about one accuracy point relative to LightGBM but reduces p95 CPU latency from 2546.06 ms to 3.26 ms. Abstaining on the lowest confidence predictions moves accepted risk down from 0.0143 to 0.00457 (SMS) and 0.1282 to 0.1006 (IMDb). Thresholds are chosen on validation/calibration data, then frozen for a single test run. More generally, the results establish calibrated lightweight classifiers as strong and transparent baselines for cost-effective low-latency text classification in real-world deployments.

GS Cited Analysis

Downloads

Download data is not yet available.

References

Alves, J. V., Leitão, D., Jesus, S., Sampaio, M. O. P., Liébana, J., Saleiro, P., Figueiredo, M. A. T., & Bizarro, P. (2025). A benchmarking framework and dataset for learning to defer in human-AI decision-making. Scientific Data, 12, Article 506. https://doi.org/10.1038/s41597-025-04664-y

Campos, M., Farinhas, A., Zerva, C., Figueiredo, M. A. T., & Martins, A. F. T. (2024). Conformal prediction for natural language processing: A survey. Transactions of the Association for Computational Linguistics, 12, 1497–1516. https://doi.org/10.1162/tacl_a_00715

Dantas, P. V., da Silva Jr., W. S., Cordeiro, L. C., & Carvalho, C. B. (2024). A comprehensive review of model compression techniques in machine learning. Applied Intelligence, 54(22), 11804–11844. https://doi.org/10.1007/s10489-024-05747-w

Deng, W., Pei, J., Ren, Z., Chen, Z., & Ren, P. (2023). Intent-calibrated self-training for answer selection in open-domain dialogues. Transactions of the Association for Computational Linguistics, 11, 1232–1249. https://doi.org/10.1162/tacl_a_00599

Dimitriadis, T., Gneiting, T., Jordan, A. I., & Vogel, P. (2024). Evaluating probabilistic classifiers: The triptych. International Journal of Forecasting, 40(3), 1101–1122. https://doi.org/10.1016/j.ijforecast.2023.09.007

Dussert, G., Chamaillé-Jammes, S., Dray, S., & Miele, V. (2024). Being confident in confidence scores: Calibration in deep learning models for camera trap image sequences. Remote Sensing in Ecology and Conservation, 11(1), 88–99. https://doi.org/10.1002/rse2.412

Fontana, M., Zeni, G., & Vantini, S. (2023). Conformal prediction: A unified review of theory and new challenges. Bernoulli, 29(1), 1–23. https://doi.org/10.3150/21-BEJ1447

Hendrickx, K., Perini, L., Van der Plas, D., Meert, W., & Davis, J. (2024). Machine learning with a reject option: A survey. Machine Learning, 113(5), 3073–3110. https://doi.org/10.1007/s10994-024-06534-x

Johansson, U., & Sönströd, C. (2026). Conformalized classifiers with reject option. Machine Learning with Applications, 23, Article 100838. https://doi.org/10.1016/j.mlwa.2026.100838

Karlsson, M., & Hössjer, O. (2024). Classification under partial reject options. Journal of Classification, 41(1), 2–37. https://doi.org/10.1007/s00357-023-09455-x

Kashani Motlagh, N., Davis, J., Anderson, T., & Gwinnup, J. (2025). Naturally constrained reject option classification. Machine Vision and Applications, 36, Article 9. https://doi.org/10.1007/s00138-024-01620-5

Kwon, H., & Kim, D.-J. (2026). Conformal selective prediction with cost aware deferral for safe clinical triage under distribution shift. Scientific Reports, 16, Article 10016. https://doi.org/10.1038/s41598-026-40637-w

Narejo, K. R., Zan, H., Mostafa, S. M., Karim, F. K., Mehmood, F., & Yaseen, A. (2026). Enhancing multimodal sentiment analysis reliability: SentiGuard+ with Dirichlet evidence and selective prediction. Journal of King Saud University Computer and Information Sciences, 38, Article 89. https://doi.org/10.1007/s44443-025-00447-y

Ojeda, F. M., Jansen, M. L., Thiéry, A., Blankenberg, S., Weimar, C., Schmid, M., & Ziegler, A. (2023). Calibrating machine learning approaches for probability estimation: A comprehensive comparison. Statistics in Medicine, 42(29), 5451–5478. https://doi.org/10.1002/sim.9921

Perez, D. M., Kaiser, B. E., & Boureima, I. (2025). Uncertainty quantification by large language models. Machine Learning with Applications, 22, Article 100773. https://doi.org/10.1016/j.mlwa.2025.100773

Silva Filho, T., Song, H., Perello-Nieto, M., Santos-Rodriguez, R., Kull, M., & Flach, P. (2023). Classifier calibration: A survey on how to assess and improve predicted class probabilities. Machine Learning, 112(9), 3211–3260. https://doi.org/10.1007/s10994-023-06336-7

Stengel-Eskin, E., & Van Durme, B. (2023). Calibrated interpretation: Confidence estimation in semantic parsing. Transactions of the Association for Computational Linguistics, 11, 1213–1231. https://doi.org/10.1162/tacl_a_00598

Steyvers, M., Tejeda, H., Kumar, A., Belem, C., Karny, S., Hu, X., Mayer, L. W., & Smyth, P. (2025). What large language models know and what people think they know. Nature Machine Intelligence, 7(2), 221–231. https://doi.org/10.1038/s42256-024-00976-7

Swaminathan, A., Lopez, I., Wang, W., Srivastava, U., Tran, E., Bhargava-Shah, A., Wu, J. Y., et al. (2024). Selective prediction for extracting unstructured clinical data. Journal of the American Medical Informatics Association, 31(1), 188–197. https://doi.org/10.1093/jamia/ocad182

Tosaki, T., Uchino, E., Kojima, R., Mineharu, Y., Okamoto, Y., Arita, M., Miyai, N., et al. (2025). Out-of-distribution reject option method for dataset shift problem in early disease onset prediction. Scientific Reports, 15, Article 19240. https://doi.org/10.1038/s41598-025-01811-8

Zhang, M., Fernández-Torres, M.-Á., Cohrs, K.-H., & Camps-Valls, G. (2025). Calibration and uncertainty quantification for deep learning-based drought detection. International Journal of Applied Earth Observation and Geoinformation, 140, Article 104563. https://doi.org/10.1016/j.jag.2025.104563

Zheng, Y., Chen, Y., Qian, B., Shi, X., Shu, Y., & Chen, J. (2025). A review on edge large language models: Design, execution, and applications. ACM Computing Surveys, 57(8), Article 209. https://doi.org/10.1145/3719664

Downloads


Crossmark Updates

How to Cite

Alhusseny, S., Alniemi, O., Saadoon, A., & Mahmood, H. (2026). Cost-Aware Selective Text Classification with Calibrated Lightweight Models. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(4), 2488-2495. https://doi.org/10.33395/sinkron.v10i4.16650