Optimizing Latency and Energy Efficiency in Edge-Native Large Language Models (LLMs) for Autonomous Mobile Agents
DOI:
https://doi.org/10.3991/ijim.v20i15.62599Keywords:
Large Language Model, Energy Efficiency, Artificial Intelligence, Inference Latency, Autonomous MobileAbstract
Large language models (LLMs) are becoming more and more important for perception, reasoning, navigation, and human-machine interaction in autonomous mobile agents like delivery robots, drones, and intelligent cars. However, the rigorous real-time latency requirements, energy limitations, and restricted processing resources make it difficult to implement LLMs directly on edge devices. Therefore, this study introduces an edge-native framework for optimizing latency and energy efficiency in LLM-enabled autonomous mobile agents. To obtain effective on-device intelligence while preserving effective performance, the suggested method combines lightweight model architectures, adaptive inference scheduling, dynamic job offloading, and hardware-aware optimization techniques. While edge–cloud collaboration allows for the selective execution of computationally demanding activities, quantization, pruning, and knowledge distillation are used to lower model complexity and memory consumption. Furthermore, an energy-aware resource management technique constantly modifies processor workloads according to task urgency, network conditions, and battery levels. When compared to traditional cloud-dependent LLM deployments, experimental evaluations on representative mobile robotic platforms show notable improvements in inference latency and battery consumption. The findings show decreased communication overhead, increased operational continuity, and faster response times without significantly lowering language comprehension or decision-making precision.
References
[1] Niu, C., Zhang, W., Zhao, Y. and Chen, Y., 2025. Energy efficient or exhaustive? Benchmarking the power consumption of llm inference engines. ACM SIGENERGY Energy Informatics Review, 5(2), pp.56-62.
[2] Wallace, T., Ombuki-Berman, B. and Ezzati-Jivan, N., 2025, May. Optimization strategies for enhancing resource efficiency in transformers & large language models. In Proceedings of the 16th ACM/SPEC International Conference on Performance Engineering (pp. 105-112).
[3] Yuan, Z., Sun, W., Liu, Y., Zhou, H., Zhou, R., Li, Y., Zhang, Z., Song, W., Huang, Y., Jia, H. and Murugesan, K., 2025. EfficientLLM: Efficiency in Large Language Models. arXiv preprint arXiv:2505.13840.
[4] Zhang, L. and Chen, Z., 2025. Opportunities of applying Large Language Models in building energy sector. Renewable and Sustainable Energy Reviews, 214, p.115558.
[5] Walkowiak, T., 2025, May. Energy Efficiency in Large Language Models: An Empirical Study. In International Conference on Dependability and Complex Systems (pp. 221-228). Cham: Springer Nature Switzerland.
[6] Zhang, L. and Chen, Z., 2023. Opportunities and challenges of applying large language models in building energy efficiency and decarbonization studies: An exploratory overview. arXiv preprint arXiv:2312.11701.
[7] Bai, G., Chai, Z., Ling, C., Wang, S., Lu, J., Zhang, N., Shi, T., Yu, Z., Zhu, M., Zhang, Y. and Song, X., 2024. Beyond efficiency: A systematic survey of resource-efficient large language models. arXiv preprint arXiv:2401.00625.
[8] Khan, T., Motie, S., Kocak, S.A. and Raza, S., 2025, May. Optimizing large language models: Metrics, energy efficiency, and case study insights. In 2025 IEEE Conference on Artificial Intelligence (CAI) (pp. 370-375). IEEE.
[9] Peng, H., Gupte, A., Eliopoulos, N.J., Ho, C.C., Mantri, R., Deng, L., Jiang, W., Lu, Y.H., Läufer, K., Thiruvathukal, G.K. and Davis, J.C., 2024. Large language models for energy-efficient code: emerging results and future directions. arXiv preprint arXiv:2410.09241.
[10] Argerich, M.F. and Patiño-Martínez, M., 2024. Measuring and improving the energy efficiency of large language models inference. IEEE Access, 12, pp.80194-80207.
[11] Pronk, K. and Zhao, Q., 2025. Benchmarking Energy Efficiency of Large Language Models Using vLLM. arXiv preprint arXiv:2509.08867.
[12] Wallace, T., Ombuki-Berman, B. and Ezzati-Jivan, N., 2025, May. Optimization strategies for enhancing resource efficiency in transformers & large language models. In Proceedings of the 16th ACM/SPEC International Conference on Performance Engineering (pp. 105-112).
[13] Kumar, M., 2023. Energy-Efficient AI Optimizing Large Language Models for Low-Power Edge Computing. International Journal of Research Publications in Engineering, Technology and Management (IJRPETM), 6(6), pp.9692-9698.
[14] Shen, Y., Shao, J., Zhang, X., Lin, Z., Pan, H., Li, D., Zhang, J. and Letaief, K.B., 2024. Large language models empowered autonomous edge AI for connected intelligence. IEEE Communications Magazine, 62(10), pp.140-146.
[15] Aliazam, M., Javadi, A., Hosseini Monazzah, A.M. and Akbari Azirani, A., 2025. RevEAL: Reliability vs Energy Optimization for Autonomous Vehicles Using Large Language Models. Scientia Iranica.c
[16] Lv, K., Huang, S., Yao, Y., Jiang, W., Huang, Y. and Feng, Z., 2025. Large Language Model-Empowered Energy-Efficient Multi-UAV-Assisted MEC Heterogeneous Networks. IEEE Transactions on Cognitive Communications and Networking.
[17] Yuan, X., Li, H., Ota, K. and Dong, M., 2024, May. Generative inference of large language models in edge computing: An energy efficient approach. In 2024 International Wireless Communications and Mobile Computing (IWCMC) (pp. 244-249). IEEE.
[18] Liang, L., Ye, H., Sheng, Y., Wang, O., Wang, J., Jin, S. and Li, G.Y., 2026. Large language models for wireless communications: From adaptation to autonomy. IEEE Communications Magazine.
[19] Sell, R., Razdan, R., Kase, K., & Rüütmann, T. (2025). The Role of AI Chatbots in Engineering Education: Experimental Findings and Implementation Strategies. International Journal of Engineering Pedagogy (iJEP), 15(5), pp. 4–19. https://doi.org/10.3991/ijep.v15i5.56681
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 A. Rajalakshmi, D. Saveetha, S. V. Manikanthan, T. Padmapriya, U. Sakthivelu

This work is licensed under a Creative Commons Attribution 4.0 International License.

