Bayesian-Optimized Word2Vec-BiLSTM for Malicious Prompt Detection in LLMs
DOI:
10.33395/sinkron.v10i4.16700Keywords:
Bayesian Optimization, BiLSTM, Deep Learning, Malicious Prompt Detection, Word2VecAbstract
Large Language Models (LLMs) are increasingly utilized across a wide range of artificial intelligence applications, but their growing adoption also introduces security concerns related to malicious prompts that may manipulate model behavior. This study investigates an optimized Word2Vec-BiLSTM approach for detecting binary malicious prompts using the Malicious Prompt Detection Dataset (MPDD). Four deep learning configurations, namely FastText-BiLSTM, FastText-BiGRU, Word2Vec-BiLSTM, and Word2Vec-BiGRU, were evaluated alongside a TF-IDF combined with Logistic Regression (LogReg) baseline under consistent experimental conditions and across three random seeds. Bayesian Optimization was applied to identify effective hyperparameter configurations and quantify performance differences between manually configured and optimized models. The models were evaluated using precision, recall, F1-score, accuracy, false-positive rate (FPR), false-negative rate (FNR), precision-recall area under the curve (PR-AUC), and computational cost. The optimized Word2Vec-BiLSTM with seed 128 achieved the highest overall performance across the primary classification metrics: 97.03% precision, 96.86% recall, 96.89% F1-score, and 96.89% accuracy, along with an FPR of 0.0070, an FNR of 0.0557, and a PR-AUC of 99.39%. The optimized deep learning models generally outperformed the TF-IDF + LogReg baseline and exhibited relatively similar computational characteristics across the evaluated deep learning configurations. These findings demonstrate that Bayesian hyperparameter optimization can improve the performance of the Word2Vec-BiLSTM model for binary malicious prompt detection within the MPDD-based experimental setting.
Downloads
References
Ahmed, S. F., Alam, Md. S. Bin, Hassan, M., Rozbu, M. R., Ishtiak, T., Rafa, N., Mofijur, M., Shawkat Ali, A. B. M., & Gandomi, A. H. (2023). Deep learning modelling techniques: current progress, applications, advantages, and challenges. Artificial Intelligence Review, 56(11), 13521–13617. https://doi.org/10.1007/s10462-023-10466-8
Ali, M. D., Saleem, A., Elahi, H., Khan, M. A., Khan, M. I., Yaqoob, M. M., Farooq Khattak, U., & Al-Rasheed, A. (2023). Breast Cancer Classification through Meta-Learning Ensemble Technique Using Convolution Neural Networks. Diagnostics, 13(13), 2242. https://doi.org/10.3390/diagnostics13132242
Alshammari, A. A., & Alsaleh, O. I. (2026). Detecting Prompt Injection Attacks in Generative AI Systems: A Hybrid SIEM and One-Class SVM Framework. Electronics, 15(11), 2242. https://doi.org/10.3390/electronics15112242
Bokolo, B. G., & Liu, Q. (2023). Deep Learning-Based Depression Detection from Social Media: Comparative Evaluation of ML and Transformer Techniques. Electronics, 12(21), 4396. https://doi.org/10.3390/electronics12214396
Bolikulov, F., Nasimov, R., Rashidov, A., Akhmedov, F., & Cho, Y.-I. (2024). Effective Methods of Categorical Data Encoding for Artificial Intelligence Algorithms. Mathematics, 12(16), 2553. https://doi.org/10.3390/math12162553
Bongirwar, V., & Mokhade, A. S. (2024). A Hybrid Bidirectional Long Short-Term Memory and Bidirectional Gated Recurrent Unit Architecture for Protein Secondary Structure Prediction. IEEE Access, 12, 115346–115355. https://doi.org/10.1109/ACCESS.2024.3444468
Cihan, P. (2025). Bayesian Hyperparameter Optimization of Machine Learning Models for Predicting Biomass Gasification Gases. Applied Sciences, 15(3), 1018. https://doi.org/10.3390/app15031018
Derner, E., Batistič, K., Zahálka, J., & Babuška, R. (2024). A Security Risk Taxonomy for Prompt-Based Interaction With Large Language Models. IEEE Access, 12, 126176–126187. https://doi.org/10.1109/ACCESS.2024.3450388
Feretzakis, G., & Verykios, V. S. (2024). Trustworthy AI: Securing Sensitive Data in Large Language Models. AI, 5(4), 2773–2800. https://doi.org/10.3390/ai5040134
Jebbar, M. A. (2025). Malicious Prompt Detection Dataset (MPDD).
Jelodar, M. B. (2025). Generative AI, Large Language Models, and ChatGPT in Construction Education, Training, and Practice. Buildings, 15(6), 933. https://doi.org/10.3390/buildings15060933
Kurniawan, A., & Chandra, M. B. (2025). Simulation of prompt injection attacks on generative pre-trained transformers models. Procedia Computer Science, 269, 400–410. https://doi.org/10.1016/j.procs.2025.08.292
Kushnerov, O., Shevchuk, R., Yevseiev, S., & Karpiński, M. (2026). Comparative Benchmarking of Deep Learning Architectures for Detecting Adversarial Attacks on Large Language Models. Information (Switzerland), 17(2). https://doi.org/10.3390/info17020155
Lan, Q., Kaul, A., & Jones, S. (2025). Prompt Injection Detection in LLM Integrated Applications. International Journal of Network Dynamics and Intelligence, 4(2). https://doi.org/10.53941/ijndi.2025.100013
Li, L., Li, J., Wang, H., & Nie, J. (2024). Application of the transformer model algorithm in chinese word sense disambiguation: a case study in chinese language. Scientific Reports, 14(1), 6320. https://doi.org/10.1038/s41598-024-56976-5
Lubis, A. R., Lase, Y. Y., Rahman, D. A., & Witarsyah, D. (2023). Improving Spell Checker Performance for Bahasa Indonesia Using Text Preprocessing Techniques with Deep Learning Models. Ingénierie Des Systèmes d Information, 28(5), 1335–1342. https://doi.org/10.18280/isi.280522
Mirshekali, H., Reza Shadi, M., Ghanadi Ladani, F., & Reza Shaker, H. (2025). A Review of Large Language Models for Energy Systems: Applications, Challenges, and Future Prospects. IEEE Access, 13, 163162–163188. https://doi.org/10.1109/ACCESS.2025.3610994
Mutinda, J., Mwangi, W., & Okeyo, G. (2023). Sentiment Analysis of Text Reviews Using Lexicon-Enhanced Bert Embedding (LeBERT) Model with Convolutional Neural Network. Applied Sciences, 13(3), 1445. https://doi.org/10.3390/app13031445
Patil, R., Boit, S., Gudivada, V., & Nandigam, J. (2023). A Survey of Text Representation and Embedding Techniques in NLP. IEEE Access, 11, 36120–36146. https://doi.org/10.1109/ACCESS.2023.3266377
Pavlatos, C., Makris, E., Fotis, G., Vita, V., & Mladenov, V. (2023). Enhancing Electrical Load Prediction Using a Bidirectional LSTM Neural Network. Electronics (Switzerland), 12(22). https://doi.org/10.3390/electronics12224652
Raza, M., Jahangir, Z., Riaz, M. B., Saeed, M. J., & Sattar, M. A. (2025). Industrial applications of large language models. Scientific Reports, 15(1), 13755. https://doi.org/10.1038/s41598-025-98483-1
Riyanto, S., Sitanggang, I. S., Djatna, T., & Atikah, T. D. (2023). Comparative Analysis using Various Performance Metrics in Imbalanced Data for Multi-class Text Classification. International Journal of Advanced Computer Science and Applications, 14(6). https://doi.org/10.14569/IJACSA.2023.01406116
Salehin, I., & Kang, D.-K. (2023). A Review on Dropout Regularization Approaches for Deep Neural Networks within the Scholarly Domain. Electronics, 12(14), 3106. https://doi.org/10.3390/electronics12143106
Shaday, E. N., Engel, V. J. L., & Heryanto, H. (2024). Application of the Bidirectional Long Short-Term Memory Method with Comparison of Word2Vec, GloVe, and FastText for Emotion Classification in Song Lyrics. Procedia Computer Science, 245, 137–146. https://doi.org/10.1016/j.procs.2024.10.237
Sujon, K. M., Hassan, R., Choi, K., & Samad, M. A. (2025). Accuracy, precision, recall, f1-score, or MCC? empirical evidence from advanced statistics, ML, and XAI for evaluating business predictive models. Journal of Big Data, 12(1), 268. https://doi.org/10.1186/s40537-025-01313-4
Varoquaux, G., & Colliot, O. (2023). Evaluating Machine Learning Models and Their Diagnostic Value. In O. Colliot (Ed.), Machine Learning for Brain Disorders (pp. 601–630). Springer US. https://doi.org/10.1007/978-1-0716-3195-9_20
Venkatesh, S., Sindhu, R., & Arunachalam, V. (2025). Hardware efficient approximate sigmoid activation function for classifying features around zero. Integration, 103, 102421. https://doi.org/10.1016/j.vlsi.2025.102421
Vyas, K. P., Bhatt, M., Dudhatra, D., Gupta, R., & Sharma, A. (2026). Malicious Prompt Classifier with Leave-One-Out Deletion Approach for Prompt Sanitization. Eighth International Conference on Futuristic Trends in Networks and Computing Technologies (FTNCT08), 3015–3024. www.sciencedirect.com
Wicaksana, H. S., Kusumaningrum, R., & Gernowo, R. (2024). Determining community happiness index with transformers and attention-based deep learning. IAES International Journal of Artificial Intelligence (IJ-AI), 13(2), 1753. https://doi.org/10.11591/ijai.v13.i2.pp1753-1761
Wicaksono, G. W., Al asqalani, S. F., Azhar, Y., Hidayah, N. P., & Andreawana, A. (2023). Automatic Summarization of Court Decision Documents over Narcotic Cases Using BERT. JOIV : International Journal on Informatics Visualization, 7(2), 416. https://doi.org/10.30630/joiv.7.2.1811
Downloads
How to Cite
Issue
Section
License
Copyright (c) 2026 Hilman Singgih Wicaksana, Gregorius Airlangga

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.






















Moraref
PKP Index
Indonesia OneSearch
OCLC Worldcat
Index Copernicus
Scilit
