An optimized CNN–LSTM model for analyzing data from surveillance cameras
Main Article Content
Abstract
This paper presents an artificial intelligence–based approach for analyzing educational processes using data obtained from surveillance cameras. The increasing deployment of surveillance cameras in urban, industrial, and transportation environments has led to the generation of massive volumes of video data, creating significant challenges for automated monitoring and real-time analysis. Traditional computer vision methods often struggle with variations in lighting, occlusions, and dynamic scenes, resulting in reduced accuracy and efficiency. This study proposes an optimized Convolutional Neural Network–Long Short-Term Memory (CNN–LSTM) model designed to address these limitations by simultaneously extracting spatial and temporal features from surveillance video streams. The proposed model is evaluated on benchmark surveillance datasets, demonstrating superior performance in object detection, activity recognition, and anomaly detection compared to conventional CNN, LSTM, and baseline hybrid approaches. Experimental results indicate that the optimized CNN–LSTM model achieves higher precision, recall, and F1-scores while maintaining real-time processing capabilities, making it suitable for practical deployment in large-scale surveillance systems. The study provides a foundation for further research in real-time surveillance analytics and intelligent video monitoring systems.
Article Details
References
Hochreiter, S., Schmidhuber, J. (1997). Long Short-Term Memory. Neural Computation, 9(8), 1735–1780. DOI: https://doi.org/10.1162/neco.1997.9.8.1735
Simonyan, K., Zisserman, A. (2015). Very Deep Convolutional Networks for Large-Scale Image Recognition. ICLR.
Goodfellow, I., Bengio, Y., Courville, A. (2016). Deep Learning. MIT Press.
Wang, H. et al. (2021). Video Behavior Recognition Using CNN-LSTM. IEEE Transactions on Image Processing.
Krizhevsky, A., Sutskever, I., Hinton, G. (2012). ImageNet Classification with Deep Convolutional Neural Networks. NIPS.
Tran, D. et al. (2015). Learning Spatiotemporal Features with 3D Convolutional Networks. ICCV. DOI: https://doi.org/10.1109/ICCV.2015.510
Sattorov, A. A. (2023). Sun’iy intellekt asosida video ma’lumotlarni tahlil qilish usullari. Toshkent: TATU nashriyoti.
Abduvaliyev, M. B. (2024). Chuqur o‘rganish texnologiyalarida vaqtli tahlil modellari. “Informatika va axborot texnologiyalari” jurnali, № 2.
Qodirov, S. & Jo‘rayev, N. (2023). Video oqimlar asosida faoliyatni aniqlash algoritmlari. TATU ilmiy axborotlari, № 4.
Rahmonov, D. (2022). Tasvir va video tahlilida neyron tarmoqlarning qo‘llanilishi. Toshkent: Fan va texnologiya.
Niyozov, M. (2024). CNN-LSTM modellarining ta’lim jarayonida kuzatuv kameralaridan foydalanishdagi roli. “Axborot tizimlari va texnologiyalari” jurnali, № 1.
Shukurov, A. (2023). Raqamli kuzatuv tizimlarida chuqur o‘rganish algoritmlari. TATU talabalari ilmiy to‘plami.
Mamatqulov, B. (2022). Sun’iy intellekt va mashinaviy o‘rganishning qo‘llanma asoslari. Toshkent: Iqtisod-Moliya.
Yusupov, S. (2023). Video ma’lumotlarni qayta ishlashda LSTM modellarining afzalliklari. “Axborot texnologiyalari” ilmiy jurnali.
Xusanov, A. & Xolmatov, D. (2024). Neyron tarmoqlar yordamida faoliyatni tasniflashning amaliy modeli. Samarqand davlat universiteti nashriyoti.
Yue-Hei Ng, J. et al. (2015). Beyond Short Snippets: Deep Networks for Video Classification. CVPR. DOI: https://doi.org/10.1109/CVPR.2015.7299101
Wang, L., Xiong, Y., Wang, Z., et al. (2016). Temporal Segment Networks for Action Recognition in Videos. ECCV.
Zhang, H., Goodfellow, I., et al. (2019). Self-Attention Mechanisms in Neural Networks for Video Understanding. arXiv preprint arXiv:1901.00059.
Qiu, Z., Yao, T., & Mei, T. (2017). Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks. ICCV. DOI: https://doi.org/10.1109/ICCV.2017.590
