Bayesian-Optimized Word2Vec-BiLSTM for Malicious Prompt Detection in LLMs

Authors

  • Hilman Singgih Wicaksana Informatics Program, Universitas Karya Husada, Semarang, Indonesia
  • Gregorius Airlangga Information Systems Program, Universitas Katolik Indonesia Atma Jaya, Jakarta, Indonesia

DOI:

10.33395/sinkron.v10i4.16700

Keywords:

Bayesian Optimization, BiLSTM, Deep Learning, Malicious Prompt Detection, Word2Vec

Abstract

Large Language Models (LLMs) are increasingly utilized across a wide range of artificial intelligence applications, but their growing adoption also introduces security concerns related to malicious prompts that may manipulate model behavior. This study investigates an optimized Word2Vec-BiLSTM approach for detecting binary malicious prompts using the Malicious Prompt Detection Dataset (MPDD). Four deep learning configurations, namely FastText-BiLSTM, FastText-BiGRU, Word2Vec-BiLSTM, and Word2Vec-BiGRU, were evaluated alongside a TF-IDF combined with Logistic Regression (LogReg) baseline under consistent experimental conditions and across three random seeds. Bayesian Optimization was applied to identify effective hyperparameter configurations and quantify performance differences between manually configured and optimized models. The models were evaluated using precision, recall, F1-score, accuracy, false-positive rate (FPR), false-negative rate (FNR), precision-recall area under the curve (PR-AUC), and computational cost. The optimized Word2Vec-BiLSTM with seed 128 achieved the highest overall performance across the primary classification metrics: 97.03% precision, 96.86% recall, 96.89% F1-score, and 96.89% accuracy, along with an FPR of 0.0070, an FNR of 0.0557, and a PR-AUC of 99.39%. The optimized deep learning models generally outperformed the TF-IDF + LogReg baseline and exhibited relatively similar computational characteristics across the evaluated deep learning configurations. These findings demonstrate that Bayesian hyperparameter optimization can improve the performance of the Word2Vec-BiLSTM model for binary malicious prompt detection within the MPDD-based experimental setting.

GS Cited Analysis

Downloads

Download data is not yet available.

References

Ahmed, S. F., Alam, Md. S. Bin, Hassan, M., Rozbu, M. R., Ishtiak, T., Rafa, N., Mofijur, M., Shawkat Ali, A. B. M., & Gandomi, A. H. (2023). Deep learning modelling techniques: current progress, applications, advantages, and challenges. Artificial Intelligence Review, 56(11), 13521–13617. https://doi.org/10.1007/s10462-023-10466-8

Ali, M. D., Saleem, A., Elahi, H., Khan, M. A., Khan, M. I., Yaqoob, M. M., Farooq Khattak, U., & Al-Rasheed, A. (2023). Breast Cancer Classification through Meta-Learning Ensemble Technique Using Convolution Neural Networks. Diagnostics, 13(13), 2242. https://doi.org/10.3390/diagnostics13132242

Alshammari, A. A., & Alsaleh, O. I. (2026). Detecting Prompt Injection Attacks in Generative AI Systems: A Hybrid SIEM and One-Class SVM Framework. Electronics, 15(11), 2242. https://doi.org/10.3390/electronics15112242

Bokolo, B. G., & Liu, Q. (2023). Deep Learning-Based Depression Detection from Social Media: Comparative Evaluation of ML and Transformer Techniques. Electronics, 12(21), 4396. https://doi.org/10.3390/electronics12214396

Bolikulov, F., Nasimov, R., Rashidov, A., Akhmedov, F., & Cho, Y.-I. (2024). Effective Methods of Categorical Data Encoding for Artificial Intelligence Algorithms. Mathematics, 12(16), 2553. https://doi.org/10.3390/math12162553

Bongirwar, V., & Mokhade, A. S. (2024). A Hybrid Bidirectional Long Short-Term Memory and Bidirectional Gated Recurrent Unit Architecture for Protein Secondary Structure Prediction. IEEE Access, 12, 115346–115355. https://doi.org/10.1109/ACCESS.2024.3444468

Cihan, P. (2025). Bayesian Hyperparameter Optimization of Machine Learning Models for Predicting Biomass Gasification Gases. Applied Sciences, 15(3), 1018. https://doi.org/10.3390/app15031018

Derner, E., Batistič, K., Zahálka, J., & Babuška, R. (2024). A Security Risk Taxonomy for Prompt-Based Interaction With Large Language Models. IEEE Access, 12, 126176–126187. https://doi.org/10.1109/ACCESS.2024.3450388

Feretzakis, G., & Verykios, V. S. (2024). Trustworthy AI: Securing Sensitive Data in Large Language Models. AI, 5(4), 2773–2800. https://doi.org/10.3390/ai5040134

Jebbar, M. A. (2025). Malicious Prompt Detection Dataset (MPDD).

Jelodar, M. B. (2025). Generative AI, Large Language Models, and ChatGPT in Construction Education, Training, and Practice. Buildings, 15(6), 933. https://doi.org/10.3390/buildings15060933

Kurniawan, A., & Chandra, M. B. (2025). Simulation of prompt injection attacks on generative pre-trained transformers models. Procedia Computer Science, 269, 400–410. https://doi.org/10.1016/j.procs.2025.08.292

Kushnerov, O., Shevchuk, R., Yevseiev, S., & Karpiński, M. (2026). Comparative Benchmarking of Deep Learning Architectures for Detecting Adversarial Attacks on Large Language Models. Information (Switzerland), 17(2). https://doi.org/10.3390/info17020155

Lan, Q., Kaul, A., & Jones, S. (2025). Prompt Injection Detection in LLM Integrated Applications. International Journal of Network Dynamics and Intelligence, 4(2). https://doi.org/10.53941/ijndi.2025.100013

Li, L., Li, J., Wang, H., & Nie, J. (2024). Application of the transformer model algorithm in chinese word sense disambiguation: a case study in chinese language. Scientific Reports, 14(1), 6320. https://doi.org/10.1038/s41598-024-56976-5

Lubis, A. R., Lase, Y. Y., Rahman, D. A., & Witarsyah, D. (2023). Improving Spell Checker Performance for Bahasa Indonesia Using Text Preprocessing Techniques with Deep Learning Models. Ingénierie Des Systèmes d Information, 28(5), 1335–1342. https://doi.org/10.18280/isi.280522

Mirshekali, H., Reza Shadi, M., Ghanadi Ladani, F., & Reza Shaker, H. (2025). A Review of Large Language Models for Energy Systems: Applications, Challenges, and Future Prospects. IEEE Access, 13, 163162–163188. https://doi.org/10.1109/ACCESS.2025.3610994

Mutinda, J., Mwangi, W., & Okeyo, G. (2023). Sentiment Analysis of Text Reviews Using Lexicon-Enhanced Bert Embedding (LeBERT) Model with Convolutional Neural Network. Applied Sciences, 13(3), 1445. https://doi.org/10.3390/app13031445

Patil, R., Boit, S., Gudivada, V., & Nandigam, J. (2023). A Survey of Text Representation and Embedding Techniques in NLP. IEEE Access, 11, 36120–36146. https://doi.org/10.1109/ACCESS.2023.3266377

Pavlatos, C., Makris, E., Fotis, G., Vita, V., & Mladenov, V. (2023). Enhancing Electrical Load Prediction Using a Bidirectional LSTM Neural Network. Electronics (Switzerland), 12(22). https://doi.org/10.3390/electronics12224652

Raza, M., Jahangir, Z., Riaz, M. B., Saeed, M. J., & Sattar, M. A. (2025). Industrial applications of large language models. Scientific Reports, 15(1), 13755. https://doi.org/10.1038/s41598-025-98483-1

Riyanto, S., Sitanggang, I. S., Djatna, T., & Atikah, T. D. (2023). Comparative Analysis using Various Performance Metrics in Imbalanced Data for Multi-class Text Classification. International Journal of Advanced Computer Science and Applications, 14(6). https://doi.org/10.14569/IJACSA.2023.01406116

Salehin, I., & Kang, D.-K. (2023). A Review on Dropout Regularization Approaches for Deep Neural Networks within the Scholarly Domain. Electronics, 12(14), 3106. https://doi.org/10.3390/electronics12143106

Shaday, E. N., Engel, V. J. L., & Heryanto, H. (2024). Application of the Bidirectional Long Short-Term Memory Method with Comparison of Word2Vec, GloVe, and FastText for Emotion Classification in Song Lyrics. Procedia Computer Science, 245, 137–146. https://doi.org/10.1016/j.procs.2024.10.237

Sujon, K. M., Hassan, R., Choi, K., & Samad, M. A. (2025). Accuracy, precision, recall, f1-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models. Journal of Big Data, 12(1), 268. https://doi.org/10.1186/s40537-025-01313-4

Varoquaux, G., & Colliot, O. (2023). Evaluating Machine Learning Models and Their Diagnostic Value. In O. Colliot (Ed.), Machine Learning for Brain Disorders (pp. 601–630). Springer US. https://doi.org/10.1007/978-1-0716-3195-9_20

Venkatesh, S., Sindhu, R., & Arunachalam, V. (2025). Hardware efficient approximate sigmoid activation function for classifying features around zero. Integration, 103, 102421. https://doi.org/10.1016/j.vlsi.2025.102421

Vyas, K. P., Bhatt, M., Dudhatra, D., Gupta, R., & Sharma, A. (2026). Malicious Prompt Classifier with Leave-One-Out Deletion Approach for Prompt Sanitization. Eighth International Conference on Futuristic Trends in Networks and Computing Technologies (FTNCT08), 3015–3024. www.sciencedirect.com

Wicaksana, H. S., Kusumaningrum, R., & Gernowo, R. (2024). Determining community happiness index with transformers and attention-based deep learning. IAES International Journal of Artificial Intelligence (IJ-AI), 13(2), 1753. https://doi.org/10.11591/ijai.v13.i2.pp1753-1761

Wicaksono, G. W., Al asqalani, S. F., Azhar, Y., Hidayah, N. P., & Andreawana, A. (2023). Automatic Summarization of Court Decision Documents over Narcotic Cases Using BERT. JOIV : International Journal on Informatics Visualization, 7(2), 416. https://doi.org/10.30630/joiv.7.2.1811

Downloads


Crossmark Updates

How to Cite

Wicaksana, H. S., & Airlangga, G. (2026). Bayesian-Optimized Word2Vec-BiLSTM for Malicious Prompt Detection in LLMs. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 10(4), 2070-2086. https://doi.org/10.33395/sinkron.v10i4.16700