Self-Attention-Based VGG16 Approach for Sign Language Gesture Recognition in Inclusive Education

Authors

  • Naoufal El-Marzouki Mohammed V University in Rabat, Rabat, Morocco https://orcid.org/0009-0003-4844-1972
  • Imane Lasri Mohammed V University in Rabat, Rabat, Morocco
  • Fatim Ezzahrae Dorhmi Mohammed V University in Rabat, Rabat, Morocco https://orcid.org/0009-0002-5992-2512
  • Anouar Riadsolh Mohammed V University in Rabat, Rabat, Morocco
  • Mourad Elbelkacemi Mohammed V University in Rabat, Rabat, Morocco

DOI:

https://doi.org/10.3991/ijoe.v22i09.60875

Keywords:

sign language recognition, self-attention, VGG16, inclusive education

Abstract


Sign language recognition (SLR) plays an important role in improving communication between deaf or hard-of-hearing individuals and the hearing community and has promising applications in inclusive education. However, conventional convolutional neural network (CNN)-based models mainly focus on local feature extraction and may fail to effectively capture global spatial dependencies, especially when recognizing visually similar static gestures. To address this limitation, this paper proposes a self-attention-based VGG16 framework for static sign language gesture recognition of alphabetic (A–Z) and numeric (0–9) hand signs. The proposed approach integrates a multi-head self-attention (MHSA) mechanism into a pretrained VGG16 architecture in order to enhance global feature representation while preserving effective local feature extraction. Experiments conducted on a 37-class static sign language dataset show that the proposed model outperforms the baseline VGG16 architecture, achieving a test accuracy of 99.87%. The obtained results confirm the effectiveness of self-attention in improving recognition performance and prediction stability for static sign language gestures. These findings suggest that the proposed framework can serve as a promising building block for intelligent assistive technologies supporting inclusive educational environments.

Author Biographies

Naoufal El-Marzouki, Mohammed V University in Rabat, Rabat, Morocco

Naoufal El-Marzouki is a PhD student at the Faculty of Sciences Rabat, University Mohammed V in Rabat, Morocco. He is a member of the Laboratory of Conception and Systems (Electronics, Signals and Informatics). He received a Master’s degree in Mathematics from the Faculty of Sciences Kenitra. His current field of research is artificial intelligence applied to education.

Imane Lasri, Mohammed V University in Rabat, Rabat, Morocco

Imane Lasri is a Professor of artificial inteligence at ENSAM, University Mohammed V in Rabat, Morocco. She is a member of the Laboratory of Conception and Systems (Electronics, Signals and Informatics). She received a Master’s degree in Big Data Engineering from the Faculty of Sciences Rabat. She received the awards of excellence of the major winners from the Mohammed V University in Rabat in 2019. Her current field of research is pattern recognition applied to higher education using deep learning algorithms. She is interested in artificial neural networks and deep learning. She is the author of many research studies published in international journals and conference proceedings.

Anouar Riadsolh, Mohammed V University in Rabat, Rabat, Morocco

Anouar Riadsolh received his PhD in Computer Science from the Faculty of Sciences Rabat (FSR), University Mohammed V in Rabat, Morocco. He is a Professor at the FSR. He is a member of the laboratory of conception and systems (electronics, signals, and Informatics), FSR. His current research interests are focused on data mining, big data, and machine learning.

Mourad Elbelkacemi, Mohammed V University in Rabat, Rabat, Morocco

Mourad Elbelkacemi receives his PhD in Computer Science. He was the dean of the Faculty of Sciences Rabat (FSR). He is a Professor at the FSR. He is a member of the laboratory of conception and systems (electronics, signals, and Informatics), FSR, University Mohammed V in Rabat, Morocco. His main research interests are focused on electronics, education, data mining, and big data.

References

[1] World Health Organization (WHO). (2021). World report on hearing. Geneva: World Health Organization. Retrieved from https://www.who.int/publications/i/item/world-report-on-hearing (Accessed: January 20, 2026). Pitts, N. B., et al. (2017). Dental car-ies. Nat. Rev. Dis. Prim., 3, pp. 17030.

[2] Mitchell, R. E., & Karchmer, M. A. (2004). Chasing the mythical ten percent: Parental hearing status of deaf and hard of hearing students in the United States. Sign Language Studies.

[3] El-Marzouki, N., Lasri, I., Riadsolh, A., & Elbelkacemi, M. (2025). Real-time sign lan-guage recognition using parallel multi-scale CNN to enhance inclusive education for deaf and hard of hearing students. Multimedia Tools and Applications, 1-20.

[4] UNESCO. (2020). Global education monitoring report 2020: Inclusion and education: All means all. Paris: UNESCO. Retrieved from https://unesdoc.unesco.org/ark:/48223/pf0000373718 (Accessed: January 20, 2026).

[5] Safeel, M., et al. (2020). A systematic review of sign language recognition techniques and approaches. (Survey paper).

[6] Kurt, A., & Erden, M. K. (2024). Investigation of the opinions of pre-service special education teachers on the use of assistive technologies in special education. Education and Information Technologies. DOI: 10.1007/s10639-023-12278-3.

[7] Lasri, I., El-Marzouki, N., Riadsolh, A., & Elbelkacemi, M. (2023). Automated Detec-tion of Dental Caries from Oral Images using Deep Convolutional Neural Net-works. International Journal of Online & Biomedical Engineering, 19(18).

[8] El-Marzouki, et al. (2024). American Sign Language recognition framework based on Convolutional Neural Networks.

[9] M. Dr, & Bhoomika S. (2025). A Real-Time Sign Language Recognition System Using MediaPipe and Random Forest Classifier Combining with Text-to-Speech Technique. In: 2025 International Conference on Artificial Intelligence and Data Engineering. IEEE. DOI: 10.1109/AIDE64228.2025.10986900.

[10] Simonyan, K., & Zisserman, A. (2015). Very Deep Convolutional Networks for Large-Scale Image Recognition. International Conference on Learning Representations (ICLR). (arXiv:1409.1556).

[11] Kingma, D. P., & Ba, J. (2015). Adam: A Method for Stochastic Optimization. Interna-tional Conference on Learning Representations (ICLR).

[12] Sign Language Gesture Images Dataset (37 classes: A–Z, 0–9, space). Kaggle. Retrieved from https://www.kaggle.com/datasets/ahmedkhanak1995/sign-language-gesture-images-dataset

[13] Aparna C, Geetha M (2020) Cnn and stacked lstm model for indian sign language recognition. Commun Comput Inf Sci 1203:126–134225–228.

[14] Shreeyash P, Prathamesh N, Nayan N (2023) Sign language recognition: a deep convo-lutional neural network approach for accurate alphabet identification. Int Res J Mod Eng Technol Sci.

Downloads

Published

2026-09-18

How to Cite

El-Marzouki, N., Lasri, I., Dorhmi, F. E., Riadsolh, A., & Elbelkacemi, M. (2026). Self-Attention-Based VGG16 Approach for Sign Language Gesture Recognition in Inclusive Education. International Journal of Online and Biomedical Engineering (iJOE), 22(09), pp. 172–185. https://doi.org/10.3991/ijoe.v22i09.60875

Issue

Section

Papers