A Deep Learning Approach for SMS Spam Detection Using Bert and Masked Language Models: A Case Study on Arabic and English

Authors

  • Musab Al-Ghadi Department of Information Technology, Faculty of Prince Al-Hussein Bin Abdallah II for Information Technology, The Hashemite University, Zarqa, Jordan.
  • Mamoon Obiedat Department of Information Technology, Faculty of Prince Al-Hussein Bin Abdallah II for Information Technology, The Hashemite University, Zarqa, Jordan.
  • Eman Omar Department of Information Technology, Faculty of Prince Al-Hussein Bin Abdallah II for Information Technology, The Hashemite University, Zarqa, Jordan.
  • Douha Al-Odeh Department of Information Technology, Faculty of Prince Al-Hussein Bin Abdallah II for Information Technology, The Hashemite University, Zarqa, Jordan.

DOI:

https://doi.org/10.56979/1102/2026/1496

Keywords:

Arabic and English SMS, Spam Detection, Phishing Attacks, BERT, MLM, Natural Language Processing

Abstract

 SMS spam poses a significant threat to the security and privacy of mobile users, causing financial fraud, phishing attacks, and diminishing trust in mobile communication systems. With multilingual and low-resource languages (like Arabic) with their linguistic complexity and dialectal variation, as well as the lack of labeled datasets for these languages, traditional spam detection methods are ineffective. This paper presents a novel multilingual SMS spam detection framework that leverages Bidirectional Encoder Representations from Transformers (BERT) with domain-adaptive Masked Language Modeling (MLM). Two models are created and tested: fine-tuned (FT) model on SMS corpus with labels and pre-trained (MLM-SMS) model on SMS corpus using MLM pretraining first followed by supervised fine-tuning. To evaluate the models, they are applied to the UCI English SMS Spam Collection dataset (5,574 messages) and a newly created Arabic dataset of 4,623 real and AI-generated SMS messages. Both models are trained using the multilingual variant of BERT, and share the same preprocessing pipeline and fixed sequence length of 128 tokens. The results demonstrate that the MLM enhanced model always works better than the baseline. It obtains 99.2% accuracy and 98.6% F1-scores on English data, whereas BERT-SMS gets 97.3% accuracy and 97.4% F1-scores. It achieves a higher accuracy of 98.8% and F1-score of 98.3% on the Arabic data compared with a baseline accuracy of 97.5% and F1-score of 96.6%. Another significant improvement in robustness is shown by the reduced number of false positives and false negatives, confirmed by the analysis of the confusion matrix. The proposed framework is scalable, language independent, and extremely effective for spam detection in real-time SMS in multilingual low resource and high resource settings.

Downloads

Published

2026-09-01

How to Cite

Musab Al-Ghadi, Mamoon Obiedat, Eman Omar, & Douha Al-Odeh. (2026). A Deep Learning Approach for SMS Spam Detection Using Bert and Masked Language Models: A Case Study on Arabic and English. Journal of Computing & Biomedical Informatics, 11(02). https://doi.org/10.56979/1102/2026/1496

Issue

Section

Articles