UAV-STNet: A Hybrid U-Net–Transformer–LSTM Architecture for Real-Time Object Detection in UAV Video Surveillance
DOI:
https://doi.org/10.29304/jqcsm.2026.18.32716Keywords:
UAV video surveillance, verification of objects in real-time, Spatio-temporal modelling, Hybrid deep learning, U-Net, Transformer and LSTMAbstract
It is difficult to perform real time object detection in Uncrewed Aerial Vehicle(UAV) based video surveillance because of the dynamic movement of camera, scale, occlusion, variation of illumination and limited power availability on-board of the computer. Purpose: In this paper of researches, the author suggests the proposed UAV Spatial-Temporal Network(UAV-STNet), which is a hybrid model of spatio-temporal deep learning model that is expected to improve the accuracy of the detection, balance of time, and real-time performance. Approaches: techniques: The proposed framework incorporates U-Net representing the system of total encoder involving extraction of multiscale space features mode, Transformer attention module which will involve global contextual modeling and Long Short-Term Memory(LSTM) network which will involve learning the short-term dependency in sequential frames. The model has been lives with and examined against a personalised UAV video dataset using a trade mark assortment of generally utilised measures of detection, including precision, recollection, mAPat 0.5 and inference velocity (FPS). Findings: UAV-STNet gave 94.5% precision, 93.1% recall and 93.2% mAP @0.5 and also 45 FPS. Better accuracy-tolerance efficient state is observed in comparative comparison with SSD, YOLOv3 and faster R-CNN which are evidently small object in motion affected scenes. Conclusions: The spatial, contextual and temporal modelling applied as a single-stop end-to-end architecture shows a powerful and computationally efficient solution of the topic of detecting real-time UAV video items, that brings benefits in the stability and reliability of intelligent surveillance systems using aerial vehicles.
Downloads
References
REFERENCES
G. Tang, J. Ni, Y. Zhao, Y. Gu, and W. Cao, “A Survey of Object Detection for UAVs Based on Deep Learning,” Remote Sens., vol. 16, no. 1, Art. no. 149, Dec. 2023, doi: 10.3390/rs16010149.
Y. Wu and K. Tong, “Research advances on deep learning-based small object detection in UAV aerial images,” Acta Aeronaut. et Astronaut. Sinica, vol. 46, no. 3, pp. 30848–30848, Feb. 2025, doi: 10.7527/S1000-6893.2024.30848.
E. Kırac and S. Özbek, “Deep Learning Based Object Detection with Unmanned Aerial Vehicle Equipped with Embedded System,” J. Aviation, vol. 8, no. 1, pp. 15–25, Feb. 2024, doi: 10.30518/jav.. 1356997.
H. Wang, Y. Li, J. Chen, and X. Gao, “Temporal Feature Aggregation for Real-Time Video Object Detection,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 5, pp. 4123–4137, May 2024, doi: 10.1109/TCSVT.2023.3312456.
P. Jiang, W. Liu, F. Wang, and R. Wei, “Hybrid U-Net Model with Visual Transformers for Enhanced Multi-Organ Medical Image Segmentation,” Information, vol. 16, no. 2, 2025, doi: 10.3390/info16020111.
Chao, M., Peng, C., Yun, L., Zhang, C., Wang, H., and Chen, Z., “A lightweight small object detection model for UAV images based on deep semantic integration,” Scientific Reports, vol. 15, no. 1, p. 31888, Aug. 2025, doi: 10.1038/s41598-025-16878-6.
X. Liu, Q. Chen, Y. Wang, and T. Huang, “Real-time UAV small object detection: SGFNet and dynamic loss optimization,” Digital Signal Processing, vol. 150, p. 105543, 2026, doi: 10.1016/j.dsp.2025.105543.
Hua, W., and Chen, Q., “A survey of small object detection based on deep learning in aerial images,” Artificial Intelligence Review, vol. 58, no. 162, pp. 1–35, 2025, doi: 10.1007/s10462-025-11150-9.
Zhang, Y., Wu, X., Li, H., and Sun, J., “Deep learning for UAV-based object detection and tracking: challenges and opportunities,” Information Fusion, vol. 102, p. 102115, 2024, doi: 10.1016/j.inffus.2023.102115.
X. B. Xu, Z. Y. Xing, M. L. Sun, P. R. Zhang, and K. H. Yang “Enhancing UAV object detection through multiscale deformable convolutions and adaptive fusion attention,” J. Supercomput., vol. 81, no. 14, pp. 1301–1327, 2025,doi: 10.1007/s11227-025-07788-5.
W. Li, Y. Zhang, and X. Sun, “Enhancing Small Object Detection in UAV Aerial Imagery through Attention-Gated Backbone and Context-Aware Fusion,” Scientific Reports, vol. 16, Art. no. 2366, Jan. 2026, doi: 10.1038/s41598-025-32074-y.
J. Tian, Q. Jin, Y. Wang, J. Yang, S. Zhang, and D. Sun, “Performance analysis of deep learning-based object detection algorithms on COCO benchmark: a comparative study,” Journal of Engineering and Applied Science, 2024, article 76, doi: 10.1186/s44147-024-00411-z.
M. Dalal and P. Mittal, “A Systematic Review of Deep Learning-Based Object Detection in Agriculture: Methods, Challenges, and Future Directions,” Computers and Materials Continua, vol. 70, no. 1, Jun. 2025, doi: 10.32604/cmc. 2025.066056.
Wanxuan Geng, Junfan Yi, Ning Li, Chen Ji, Yu Cong, and Liang Cheng, “RCSD-UAV: An object detection dataset for unmanned aerial vehicles in realistic complex scenarios,” Engineering Applications of Artificial Intelligence, vol. 151, p. 110748, Jul. 2025, doi: 10.1016/j.engappai.2025.110748.
A. Aybilge Murat and M. Servet Kiran, “A comprehensive review on YOLO versions for object detection,” Journal of Engineering Science and Technology Review, vol. 16, 2025, doi: 10.1016/j.jestch.2025.102161.
A. Javed Sayyad and K. Attarde, “A systematic literature review on deep learning approaches for small object detection,” Array, vol. 100615, 2 Dec. 2025, doi: 10.1016/j.array.2025.100615.
S. E. Prasetyo, C. Atmaja, M. Ardian, A. Ardhiansyah, A. R. Sudarni, and M. Khaira, “Benchmarking YOLOv3 and SSD: A Performance Comparison for Multi-Object Detection,” Edu Komputika Journal, vol. 11, no. 2, pp. 136–146, 2024, doi: 10.15294/edukom.v11i2.28005.
Z. Zheng, P. Wang, W. Liu, J. Li, R. Ye, and D. Ren, “Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression,” Proceedings of AAAI, 2020, doi: 10.1609/aaai.v34i07.6999.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Shoohi, Liqaa M.

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.








