An Efficient Multimodal Causal Attention Architecture for Real-Time Mobile Educational Content Interpretation
DOI:
https://doi.org/10.3991/ijim.v20i19.63224Keywords:
Multimodal Educational Content, Causal Attention, T-DAG, Cross-Modal Layer Fusion, Causal Context-Aware QuantizationAbstract
The increasing adoption of mobile learning requires efficient multimodal models capable of real-time educational content analysis under resource constraints. However, existing multimodal architectures often involve high inference latency and memory requirements, limiting their deployment on mobile and edge devices. This study proposes the Efficient Multimodal Causal Attention Architecture (EMCAA), integrating speech, lecture transcripts, presentation slides, and contextual learner interactions. EMCAA combines Temporal Directed Acyclic Graph (T-DAG) Causal Attention to prioritize temporally relevant dependencies, Cross-Modal Layer Fusion (CMLF) to integrate complementary multimodal representations, and Causal ContextAware Quantization (CCQ) to reduce computational and memory requirements through adaptive precision allocation. Experiments on the Multimodal Student Attention Dataset showed that EMCAA achieved 91.4% accuracy, 91.0% precision, 90.8% recall, and 90.9% F1-score. Compared with MC-MIFA, EMCAA improved accuracy by 1.2 percentage points while reducing inference latency from 57 to 43 ms (24.6%) and memory consumption from 401 to 329 MB (18.0%). Ablation analysis further demonstrated the contribution of the proposed components. These findings indicate the potential of EMCAA for efficient multimodal educational content processing in resource-constrained mobile and edge environments.
References
[1] I. S. Almuniri et al., “Beyond peak accuracy: A stability-centric framework for reliable multimodal student engagement assessment,” Scientific Reports, vol. 16, no. 1, Art. no. 5, 2026.
[2] P. Bhardwaj et al., “Application of deep learning on student engagement in e-learning environments,” Computers & Electrical Engineering, vol. 93, Art. no. 107277, 2021.
[3] I. Boutabia et al., “Hybrid CNN-ViT model for student engagement detection in open classroom environments,” SN Computer Science, vol. 6, no. 6, Art. no. 684, 2025.
[4] B. D. Bhavani, R. Shetty, and D. Mahesh, “Enhancing student engagement and personalized learning through AI tools: A comprehensive review,” Computer Science and Engineering International Journal, vol. 15, no. 1, pp. 111–130, 2025, doi: 10.5121/cseij.2025.15113.
[5] G. Chen et al., “Video Mamba suite: State space model as a versatile alternative for video understanding,” International Journal of Computer Vision, vol. 134, no. 1, Art. no. 20, 2026.
[6] Z. Dan, Z. Yali, and Z. Tong, “A multimodal AI framework for real-time student engagement detection and adaptive feedback in higher education,” Journal of King Saud University—Computer and Information Sciences, vol. 38, Art. no. 243, 2026, doi: 10.1007/s44443-026-00639-0.
[7] A. A. Fraiden, “Anticipatory thinking and AI-driven assessments: A balanced approach to AI integration in education aligned with Saudi Vision 2030,” African Journal of Biomedical Research, vol. 27, no. 3, pp. 619–630, 2024, doi: 10.53555/ajbr.v27i3.2560.
[8] J. D. T. Guerrero-Sosa, F. P. Romero, V. H. M. Domínguez, J. Serrano-Guerrero, A. Montoro-Montarroso, and J. Á. Olivas, “A comprehensive review of multimodal analysis in education,” Applied Sciences, vol. 15, no. 11, Art. no. 5896, 2025, doi: 10.3390/app15115896.
[9] S. Gupta, P. Kumar, and R. Tekchandani, “A multimodal facial cues based engagement detection system in e-learning context using deep learning approach,” Multimedia Tools and Applications, vol. 82, no. 18, pp. 28589–28615, 2023, doi: 10.1007/s11042-023-14392-3.
[10] Z. Ji and S. Li, “Multimodal alignment and attention-based person search via natural language description,” IEEE Internet of Things Journal, vol. 7, no. 11, pp. 11147–11156, 2020.
[11] W. Li, F. Deng, and Z. Li, “Question-guided attention and cross-modal alignment for knowledge-based visual question answering,” Information Processing & Management, vol. 63, no. 3, Art. no. 104578, 2026.
[12] N. K. Mehta, S. S. Prasad, S. Saurav, R. Saini, and S. Singh, “Three-dimensional DenseNet self-attention neural network for automatic detection of student’s engagement,” Applied Intelligence, vol. 52, no. 12, pp. 13803–13823, 2022, doi: 10.1007/s10489-022-03200-4.
[13] Chen, H., & Shang, Y. (2026). Construction and Application of a Learning Resource Sharing Platform in Higher Education Based on Mobile Interactive Technology. International Journal of Interactive Mobile Technologies (iJIM), 20(01), pp. 19–33. https://doi.org/10.3991/ijim.v20i01.59785.
[14] F. Naseer and S. Khawaja, “Mitigating conceptual learning gaps in mixed-ability classrooms: A learning analytics-based evaluation of AI-driven adaptive feedback for struggling learners,” Applied Sciences, vol. 15, no. 8, Art. no. 4473, 2025, doi: 10.3390/app15084473.
[15] P. Thottempudi, V. Kumar, and R. Kumar, “Dynamic multi-modal attention network for robust and real-time through-wall human activity recognition,” Results in Engineering, vol. 28, Art. no. 107632, 2025, doi: 10.1016/j.rineng.2025.107632.
[16] W. Qu, L. Guo, J. Cui, and X. Jin, “Multimodal attention-based instruction-following part-level affordance grounding,” Applied Sciences, vol. 14, no. 11, Art. no. 4696, 2024, doi: 10.3390/app14114696.
[17] F. A. Raza, A. D. Singh, J. J. S. Kovilpillai, A. Hamdan, and V. Rajaratnam, “Safeguarding integrity in AI-enhanced education: Stakeholder perspectives on accuracy, validity, and ethics in ASEAN,” European Journal of STEM Education, vol. 10, no. 1, Art. no. 22, 2025, doi: 10.20897/ejsteme/17307.
[18] M. Roshanaei, H. Olivares, and R. R. Lopez, “Harnessing AI to foster equity in education: Opportunities, challenges, and emerging strategies,” Journal of Intelligent Learning Systems and Applications, vol. 15, no. 4, pp. 123–140, 2023, doi: 10.4236/jilsa.2023.154009.
[19] F. M. Shiri, E. Ahmadi, M. Rezaee, and T. Perumal, “Detection of student engagement in e-learning environments using EfficientNetV2-L together with RNN-based models,” Journal of Artificial Intelligence, vol. 6, no. 1, pp. 85–103, 2024, doi: 10.32604/jai.2024.048911.
[20] C. Wang, S. Zhang, T. Li et al., “MC-MIFA: A causal-aware hybrid state space framework for robust multimodal student engagement analysis,” Scientific Reports, vol. 16, Art. no. 20467, 2026, doi: 10.1038/s41598-026-51563-2.
[21] C. W. Woo et al., “Students’ perception of the classroom environment: A comparison between innovative and traditional classrooms,” Journal of School Teaching and Learning, vol. 22, no. 1, pp. 31–17, 2022.
[22] Wang, C. (2026). Mobile AIGC Image Generation Interactive Training Model for Design Education. International Journal of Interactive Mobile Technologies (iJIM), 20(02), pp. 63–77. https://doi.org/10.3991/ijim.v20i02.60137.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Jin Wang, Inam Ullah

This work is licensed under a Creative Commons Attribution 4.0 International License.

