Human Activity Detection and Action Recognition in Videos Using Convolutional Neural Networks
DOI:
https://doi.org/10.32890/jict2020.19.2.6Keywords:
Action recognition, convolutional neural network, Gaussian Mixture Model, optical flow, SIFT feature extractionAbstract
Human activity recognition from video scenes has become a significant area of research in the field of computer vision applications. Action recognition is one of the most challenging problems in the area of video analysis and it finds applications in human-computer interaction, anomalous activity detection, crowd monitoring and patient monitoring. Several approaches have been presented for human activity recognition using machine learning techniques. The main aim of this work is to detect and track human activity, and classify actions for two publicly available video databases. In this work, a novel approach of feature extraction from video sequence by combining Scale Invariant Feature Transform and optical flow computation are used where shape, gradient and orientation features are also incorporated for robust feature formulation. Tracking of human activity in the video is implemented using the Gaussian Mixture Model. Convolutional Neural Network based classification approach is used for database training and testing purposes. The activity recognition performance is evaluated for two public datasets namely Weizmann dataset and Kungliga Tekniska Hogskolan dataset with action recognition accuracy of 98.43% and 94.96%, respectively. Experimental and comparative studies have shown that the proposed approach outperformed state-of the art techniques.
References
Aggarwal, J. K., & Ryoo, M. S. (2011). Human activity analysis. ACM computing surveys, 43(3), 1–43. https://doi.org/10.1145/1922649.1922653
Burghouts, G. J., & Schutte, K. (2013). Spatio-temporal layout of human actions for improved bag-of-words action detection. Pattern Recognition Letters, 34(15), 1861–1869. https://doi.org/10.1016/j.patrec.2013.01.024
Cai, J., Tang, X., Zhang, L., & Feng, G. (2017). Learning zeroth class dictionary for human action recognition. Communications in Computer and Information Science, vol. 773, 651–666. https://doi.org/10.1007/978-981-10-7305-2_55
Chaaraoui, A. A., Climent-Pérez, P., & Flórez-Revuelta, F. (2013). Silhouette-based human action recognition using sequences of key poses. Pattern Journal of ICT, 19, No. 2 (April) 2020, pp: 157-Recognition Letters, 34(15), 1799–1807. https://doi.org/10.1016/j. patrec.2013.01.021
Cheema, S., Eweiwi, A., Thurau, C., & Bauckhage, C. (2011). Action recognition by learning discriminative key poses. Paper presented at the IEEE International Conference on Computer Vision Workshops (ICCV Workshops). Retrieved from https://doi.org/10.1109/ iccvw.2011.6130402
Cheng, J., Liu, H., Wang, F., Li, H., & Zhu, C. (2015). Silhouette analysis for human action recognition based on supervised temporal t-SNE and incremental learning. Paper presented at the IEEE Transactions on Image Processing, 24(10), 3203–3217. https://doi.org/10.1109/ tip.2015.2441634
Dollar, P., Rabaud, V., Cottrell, G., & Belongie, S. (2019). Behavior recognition via sparse spatio-temporal features. Paper presented at the IEEE International Workshop on Visual Surveillance and Performance Evaluation of Tracking and Surveillance. https://doi.org/10.1109/ vspets.2005.1570899
Farhadi, A., & Tabrizi, M. K. (2008). Learning to recognize activities from the wrong view point. Lecture Notes in Computer Science, vol. 5302, 154–166. https://doi.org/10.1007/978-3-540-88682-2_13
Gorelick, L., Blank, M., Shechtman, E., Irani, M., & Basri, R. (2005). Actions as space-time shapes. Paper presented at the IEEE Transactions on Pattern Analysis and Machine Intelligence, 29(12), 2247–2253. https://doi.org/10.1109/tpami.2007.70711
Ji, S., Xu, W., Yang, M., & Yu, K. (2013). 3D convolutional neural networks for human action recognition. Paper presented at the IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(1), 221–231. https://doi.org/10.1109/tpami.2012.59
Ji, X.-F., Wu, Q.-Q., Ju, Z.-J., & Wang, Y.-Y. (2014). Study of human action recognition based on improved spatio-temporal features. International Journal of Automation and Computing, 11(5), 500–509. https://doi.org/10.1007/s11633-014-0831-4
Jiang, H., Drew, M. S., & Ze-Nian Li. (2010). Action detection in cluttered video with successive convex matching. Paper presented at the IEEE Transactions on Circuits and Systems for Video Technology, 20(1), 50–64. https://doi.org/10.1109/tcsvt.2009.2026947
Kaminski, L., Mackowiak, S., & Domanski, M. (2017). Human activity recognition using standard descriptors of MPEG CDVS. Paper presented at the IEEE International Conference on Multimedia & Expo Workshops (ICMEW). https://doi.org/10.1109/icmew.2017.8026248
Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., & Fei-Fei, L. (2014). Large-scale video classification with convolutional neural Journal of ICT, 19, No. 2 (April) 2020, pp: 157-networks. Paper presented at the IEEE Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/cvpr.2014.223
Klaser, A., Marszałek, M., & Schmid, C. (2008). A spatio-temporal descriptor based on 3D-gradients. Proceedings of the British Machine Vision Conference, 99.1-99.10. https://doi.org/10.5244/C.22.99
Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2017). ImageNet classification with deep convolutional neural networks. Communications of the ACM, 60(6), 84–90. https://doi.org/10.1145/3065386
Laptev, I., Marszalek, M., Schmid, C., & Rozenfeld, B. (2008). Learning realistic human actions from movies. Paper presented at the IEEE Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/cvpr.2008.4587756
Lee, S., Le, H. X., Ngo, H. Q., Kim, H. I., Han, M., & Lee, Y. K. (2011). Semi-Markov conditional random fields for accelerometer-based activity recognition. Applied Intelligence, 35(2), 226-241.
Li, X., Zhang, Y., & Liao, D. (2017). Mining Key Skeleton Poses with Latent SVM for Action Recognition. Applied Computational Intelligence and Soft Computing, 1–11. https://doi.org/10.1155/2017/5861435
Liu, M., Chen, C., & Liu, H. (2017). Time-ordered spatial-temporal interest points for human action classification. Paper presented at the IEEE International Conference on Multimedia and Expo (ICME), 655–660. https://doi.org/10.1109/icme.2017.8019477
Lu, M., & Zhang, L. (2014). Action recognition by fusing spatial-temporal appearance and the local distribution of interest points. Proceedings of the 2014 International Conference on Future Computer and Communication Engineering. https://doi.org/10.2991/icfcce-14.2014.19
Meng, B., Liu, X., & Wang, X. (2018). Human action recognition based on quaternion spatial-temporal convolutional neural network and LSTM in RGB videos. Multimedia Tools and Applications, 77(20), 26901–26918. https://doi.org/10.1007/s11042-018-5893-9
Moghaddam, Z., & Piccardi, M. (2014). Training initialization of hidden Markov models in human action recognition. Paper presented at the IEEE Transactions on Automation Science and Engineering, 11(2), 394–408. https://doi.org/10.1109/tase.2013.2262940
Niebles, J. C., Wang, H., & Fei-Fei, L. (2008). Unsupervised learning of human action categories using spatial-temporal words. International Journal of Computer Vision, 79(3), 299–318. https://doi.org/10.1007/ s11263-007-0122-4
Ramasso, E., Panagiotakis, C., Pellerin, D., & Rombaut, M. (2007). Human action recognition in videos based on the transferable belief model. Pattern Analysis and Applications, 11(1), 1–19. https://doi.org/10.1007/ s10044-007-0073-y Journal of ICT, 19, No. 2 (April) 2020, pp: 157-
Schuldt, C., Laptev, I., & Caputo, B. (2004). Recognizing human actions: A local SVM approach. Proceedings of the 17th International Conference on Pattern Recognition, Vol.3, 32–36. https://doi.org/10.1109/ icpr.2004.1334462
Shao, L., Zhen, X., Tao, D., & Li, X. (2014). Spatio-temporal laplacian pyramid coding for action recognition. Paper presented at the IEEE Transactions on Cybernetics, 44(6), 817–827. https://doi.org/10.1109/ tcyb.2013.2273174
Taylor, G. W., Fergus, R., LeCun, Y., & Bregler, C. (2010). Convolutional Learning of Spatio-temporal Features. Computer Vision – ECCV, 140–153. https://doi.org/10.1007/978-3-642-15567-3_11
Varol, G., Laptev, I., & Schmid, C. (2018). Long-term temporal convolutions for action recognition. Paper presented at the IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(6), 1510–1517. https://doi.org/10.1109/tpami.2017.2712608
Wang, Haoran, Yuan, C., Hu, W., & Sun, C. (2012). Supervised class-specific dictionary learning for sparse modeling in action recognition. Pattern Recognition, 45(11), 3902–3911. https://doi.org/10.1016/j. patcog.2012.04.024
Wang, Heng, Kläser, A., Schmid, C., & Liu, C.-L. (2013). Dense trajectories and motion boundary descriptors for action recognition. International Journal of Computer Vision, 103(1), 60–79. https://doi.org/10.1007/ s11263-012-0594-8
Wang, Lei, Xu, Y., Cheng, J., Xia, H., Yin, J., & Wu, J. (2018). Human action recognition by learning spatio-temporal features with deep neural networks. IEEE Access, 6, 17913–17922. https://doi.org/10.1109/ access.2018.2817253
Wang, Lei, Xu, Y., Cheng, J., Xia, H., Yin, J., & Wu, J. (2018). Human action recognition by learning spatio-temporal features with deep neural networks. IEEE Access, 6, 17913–17922. https://doi.org/10.1109/ access.2018.2817253
Wang, Limin, Qiao, Y., & Tang, X. (2015). Action recognition with trajectory-pooled deep-convolutional descriptors. Paper presented at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4305–4314. https://doi.org/10.1109/cvpr.2015.7299059
Wang, Limin, Qiao, Y., & Tang, X. (2015). Action recognition with trajectory-pooled deep-convolutional descriptors. Paper presented at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4305–4314. https://doi.org/10.1109/cvpr.2015.7299059
Wong, S.-F., & Cipolla, R. (2007). Extracting spatiotemporal interest points using global information. Paper presented at the IEEE 11th International Conference on Computer Vision, 1–8. https://doi.org/10.1109/iccv.2007.4408923 Journal of ICT, 19, No. 2 (April) 2020, pp: 157-
Wong, S.-F., & Cipolla, R. (2007). Extracting spatiotemporal interest points using global information. Paper presented at the IEEE 11th International Conference on Computer Vision, 1–8. https://doi.org/10.1109/iccv.2007.4408923
Wu, S., Oreifej, O., & Shah, M. (2011). Action recognition in videos acquired by a moving camera using motion decomposition of Lagrangian particle trajectories. Paper presented at the International Conference on Computer Vision, 1419-1426. https://doi.org/10.1109/ iccv.2011.6126397
Wu, S., Oreifej, O., & Shah, M. (2011). Action recognition in videos acquired by a moving camera using motion decomposition of Lagrangian particle trajectories. International Conference on Computer Vision, 1419-1426. https://doi.org/10.1109/iccv.2011.6126397
Yao, L., Liu, Y., & Huang, S. (2016). Spatio-temporal information for human action recognition. EURASIP Journal on Image and Video Processing, 2016(1). https://doi.org/10.1186/s13640-016-0145-2
Yao, L., Liu, Y., & Huang, S. (2016). Spatio-temporal information for human action recognition. EURASIP Journal on Image and Video Processing, 2016(1). https://doi.org/10.1186/s13640-016-0145-2
Zhang, B., Yang, Y., Chen, C., Yang, L., Han, J., & Shao, L. (2017). Action recognition using 3D histograms of texture and a multi-class boosting classifier. Paper presented at the IEEE Transactions on Image Processing, 26(10), 4648–4660. https://doi.org/10.1109/ tip.2017.2718189
Zhang, B., Yang, Y., Chen, C., Yang, L., Han, J., & Shao, L. (2017). Action recognition using 3D histograms of texture and a multi-class boosting classifier. Paper presented at the IEEE Transactions on Image Processing, 26(10), 4648–4660. https://doi.org/10.1109/ tip.2017.2718189
Zhen, X., Zheng, F., Shao, L., Cao, X., & Xu, D. (2017). Supervised local descriptor learning for human action recognition. Paper presented at the IEEE Transactions on Multimedia, 19(9), 2056–2065. https://doi.org/10.1109/tmm.2017.2700204
Zhen, X., Zheng, F., Shao, L., Cao, X., & Xu, D. (2017). Supervised local descriptor learning for human action recognition. Paper presented at the IEEE Transactions on Multimedia, 19(9), 2056–2065. https://doi.org/10.1109/tmm.2017.2700204
Zhu, S., & Xia, L. (2015). Human action recognition based on fusion features extraction of adaptive background subtraction and optical flow model. Mathematical Problems in Engineering, 1–11. https://doi.org/10.1155/2015/387464
Published
Issue
Section
How to Cite
Research impact
Harvested 2026-09-25Counts differ between services because each indexes a different body of literature. None of them is the whole picture.
2002 - 2020






















