Human Activity Detection and Action Recognition in Videos Using Convolutional Neural Networks

Authors

  • Jagadeesh Basavaiah Department of Electronics and Communication Engineering, Vidyavardhaka College of Engineering, India
  • Chandrashekar Mohan Patil Department of Electronics and Communication Engineering, Vidyavardhaka College of Engineering, India

DOI:

https://doi.org/10.32890/jict2020.19.2.6

Keywords:

Action recognition, convolutional neural network, Gaussian Mixture Model, optical flow, SIFT feature extraction

Abstract

Human activity recognition from video scenes has become a significant area of research in the field of computer vision applications. Action recognition is one of the most challenging problems in the area of video analysis and it finds applications in human-computer interaction, anomalous activity detection, crowd monitoring and patient monitoring. Several approaches have been presented for human activity recognition using machine learning techniques. The main aim of this work is to detect and track human activity, and classify actions for two publicly available video databases. In this work, a novel approach of feature extraction from video sequence by combining Scale Invariant Feature Transform and optical flow computation are used where shape, gradient and orientation features are also incorporated for robust feature formulation. Tracking of human activity in the video is implemented using the Gaussian Mixture Model. Convolutional Neural Network based classification approach is used for database training and testing purposes. The activity recognition performance is evaluated for two public datasets namely Weizmann dataset and Kungliga Tekniska Hogskolan dataset with action recognition accuracy of 98.43% and 94.96%, respectively. Experimental and comparative studies have shown that the proposed approach outperformed state-of the art techniques.

References

Aggarwal, J. K., & Ryoo, M. S. (2011). Human activity analysis. ACM computing surveys, 43(3), 1–43. https://doi.org/10.1145/1922649.1922653

Burghouts, G. J., & Schutte, K. (2013). Spatio-temporal layout of human actions for improved bag-of-words action detection. Pattern Recognition Letters, 34(15), 1861–1869. https://doi.org/10.1016/j.patrec.2013.01.024

Cai, J., Tang, X., Zhang, L., & Feng, G. (2017). Learning zeroth class dictionary for human action recognition. Communications in Computer and Information Science, vol. 773, 651–666. https://doi.org/10.1007/978-981-10-7305-2_55

Chaaraoui, A. A., Climent-Pérez, P., & Flórez-Revuelta, F. (2013). Silhouette-based human action recognition using sequences of key poses. Pattern Journal of ICT, 19, No. 2 (April) 2020, pp: 157-Recognition Letters, 34(15), 1799–1807. https://doi.org/10.1016/j. patrec.2013.01.021

Cheema, S., Eweiwi, A., Thurau, C., & Bauckhage, C. (2011). Action recognition by learning discriminative key poses. Paper presented at the IEEE International Conference on Computer Vision Workshops (ICCV Workshops). Retrieved from https://doi.org/10.1109/ iccvw.2011.6130402

Cheng, J., Liu, H., Wang, F., Li, H., & Zhu, C. (2015). Silhouette analysis for human action recognition based on supervised temporal t-SNE and incremental learning. Paper presented at the IEEE Transactions on Image Processing, 24(10), 3203–3217. https://doi.org/10.1109/ tip.2015.2441634

Dollar, P., Rabaud, V., Cottrell, G., & Belongie, S. (2019). Behavior recognition via sparse spatio-temporal features. Paper presented at the IEEE International Workshop on Visual Surveillance and Performance Evaluation of Tracking and Surveillance. https://doi.org/10.1109/ vspets.2005.1570899

Farhadi, A., & Tabrizi, M. K. (2008). Learning to recognize activities from the wrong view point. Lecture Notes in Computer Science, vol. 5302, 154–166. https://doi.org/10.1007/978-3-540-88682-2_13

Gorelick, L., Blank, M., Shechtman, E., Irani, M., & Basri, R. (2005). Actions as space-time shapes. Paper presented at the IEEE Transactions on Pattern Analysis and Machine Intelligence, 29(12), 2247–2253. https://doi.org/10.1109/tpami.2007.70711

Ji, S., Xu, W., Yang, M., & Yu, K. (2013). 3D convolutional neural networks for human action recognition. Paper presented at the IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(1), 221–231. https://doi.org/10.1109/tpami.2012.59

Ji, X.-F., Wu, Q.-Q., Ju, Z.-J., & Wang, Y.-Y. (2014). Study of human action recognition based on improved spatio-temporal features. International Journal of Automation and Computing, 11(5), 500–509. https://doi.org/10.1007/s11633-014-0831-4

Jiang, H., Drew, M. S., & Ze-Nian Li. (2010). Action detection in cluttered video with successive convex matching. Paper presented at the IEEE Transactions on Circuits and Systems for Video Technology, 20(1), 50–64. https://doi.org/10.1109/tcsvt.2009.2026947

Kaminski, L., Mackowiak, S., & Domanski, M. (2017). Human activity recognition using standard descriptors of MPEG CDVS. Paper presented at the IEEE International Conference on Multimedia & Expo Workshops (ICMEW). https://doi.org/10.1109/icmew.2017.8026248

Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., & Fei-Fei, L. (2014). Large-scale video classification with convolutional neural Journal of ICT, 19, No. 2 (April) 2020, pp: 157-networks. Paper presented at the IEEE Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/cvpr.2014.223

Klaser, A., Marszałek, M., & Schmid, C. (2008). A spatio-temporal descriptor based on 3D-gradients. Proceedings of the British Machine Vision Conference, 99.1-99.10. https://doi.org/10.5244/C.22.99

Krizhevsky, A., Sutskever, I., & Hinton, G. E. (2017). ImageNet classification with deep convolutional neural networks. Communications of the ACM, 60(6), 84–90. https://doi.org/10.1145/3065386

Laptev, I., Marszalek, M., Schmid, C., & Rozenfeld, B. (2008). Learning realistic human actions from movies. Paper presented at the IEEE Conference on Computer Vision and Pattern Recognition. https://doi.org/10.1109/cvpr.2008.4587756

Lee, S., Le, H. X., Ngo, H. Q., Kim, H. I., Han, M., & Lee, Y. K. (2011). Semi-Markov conditional random fields for accelerometer-based activity recognition. Applied Intelligence, 35(2), 226-241.

Li, X., Zhang, Y., & Liao, D. (2017). Mining Key Skeleton Poses with Latent SVM for Action Recognition. Applied Computational Intelligence and Soft Computing, 1–11. https://doi.org/10.1155/2017/5861435

Liu, M., Chen, C., & Liu, H. (2017). Time-ordered spatial-temporal interest points for human action classification. Paper presented at the IEEE International Conference on Multimedia and Expo (ICME), 655–660. https://doi.org/10.1109/icme.2017.8019477

Lu, M., & Zhang, L. (2014). Action recognition by fusing spatial-temporal appearance and the local distribution of interest points. Proceedings of the 2014 International Conference on Future Computer and Communication Engineering. https://doi.org/10.2991/icfcce-14.2014.19

Meng, B., Liu, X., & Wang, X. (2018). Human action recognition based on quaternion spatial-temporal convolutional neural network and LSTM in RGB videos. Multimedia Tools and Applications, 77(20), 26901–26918. https://doi.org/10.1007/s11042-018-5893-9

Moghaddam, Z., & Piccardi, M. (2014). Training initialization of hidden Markov models in human action recognition. Paper presented at the IEEE Transactions on Automation Science and Engineering, 11(2), 394–408. https://doi.org/10.1109/tase.2013.2262940

Niebles, J. C., Wang, H., & Fei-Fei, L. (2008). Unsupervised learning of human action categories using spatial-temporal words. International Journal of Computer Vision, 79(3), 299–318. https://doi.org/10.1007/ s11263-007-0122-4

Ramasso, E., Panagiotakis, C., Pellerin, D., & Rombaut, M. (2007). Human action recognition in videos based on the transferable belief model. Pattern Analysis and Applications, 11(1), 1–19. https://doi.org/10.1007/ s10044-007-0073-y Journal of ICT, 19, No. 2 (April) 2020, pp: 157-

Schuldt, C., Laptev, I., & Caputo, B. (2004). Recognizing human actions: A local SVM approach. Proceedings of the 17th International Conference on Pattern Recognition, Vol.3, 32–36. https://doi.org/10.1109/ icpr.2004.1334462

Shao, L., Zhen, X., Tao, D., & Li, X. (2014). Spatio-temporal laplacian pyramid coding for action recognition. Paper presented at the IEEE Transactions on Cybernetics, 44(6), 817–827. https://doi.org/10.1109/ tcyb.2013.2273174

Taylor, G. W., Fergus, R., LeCun, Y., & Bregler, C. (2010). Convolutional Learning of Spatio-temporal Features. Computer Vision – ECCV, 140–153. https://doi.org/10.1007/978-3-642-15567-3_11

Varol, G., Laptev, I., & Schmid, C. (2018). Long-term temporal convolutions for action recognition. Paper presented at the IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(6), 1510–1517. https://doi.org/10.1109/tpami.2017.2712608

Wang, Haoran, Yuan, C., Hu, W., & Sun, C. (2012). Supervised class-specific dictionary learning for sparse modeling in action recognition. Pattern Recognition, 45(11), 3902–3911. https://doi.org/10.1016/j. patcog.2012.04.024

Wang, Heng, Kläser, A., Schmid, C., & Liu, C.-L. (2013). Dense trajectories and motion boundary descriptors for action recognition. International Journal of Computer Vision, 103(1), 60–79. https://doi.org/10.1007/ s11263-012-0594-8

Wang, Lei, Xu, Y., Cheng, J., Xia, H., Yin, J., & Wu, J. (2018). Human action recognition by learning spatio-temporal features with deep neural networks. IEEE Access, 6, 17913–17922. https://doi.org/10.1109/ access.2018.2817253

Wang, Lei, Xu, Y., Cheng, J., Xia, H., Yin, J., & Wu, J. (2018). Human action recognition by learning spatio-temporal features with deep neural networks. IEEE Access, 6, 17913–17922. https://doi.org/10.1109/ access.2018.2817253

Wang, Limin, Qiao, Y., & Tang, X. (2015). Action recognition with trajectory-pooled deep-convolutional descriptors. Paper presented at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4305–4314. https://doi.org/10.1109/cvpr.2015.7299059

Wang, Limin, Qiao, Y., & Tang, X. (2015). Action recognition with trajectory-pooled deep-convolutional descriptors. Paper presented at the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4305–4314. https://doi.org/10.1109/cvpr.2015.7299059

Wong, S.-F., & Cipolla, R. (2007). Extracting spatiotemporal interest points using global information. Paper presented at the IEEE 11th International Conference on Computer Vision, 1–8. https://doi.org/10.1109/iccv.2007.4408923 Journal of ICT, 19, No. 2 (April) 2020, pp: 157-

Wong, S.-F., & Cipolla, R. (2007). Extracting spatiotemporal interest points using global information. Paper presented at the IEEE 11th International Conference on Computer Vision, 1–8. https://doi.org/10.1109/iccv.2007.4408923

Wu, S., Oreifej, O., & Shah, M. (2011). Action recognition in videos acquired by a moving camera using motion decomposition of Lagrangian particle trajectories. Paper presented at the International Conference on Computer Vision, 1419-1426. https://doi.org/10.1109/ iccv.2011.6126397

Wu, S., Oreifej, O., & Shah, M. (2011). Action recognition in videos acquired by a moving camera using motion decomposition of Lagrangian particle trajectories. International Conference on Computer Vision, 1419-1426. https://doi.org/10.1109/iccv.2011.6126397

Yao, L., Liu, Y., & Huang, S. (2016). Spatio-temporal information for human action recognition. EURASIP Journal on Image and Video Processing, 2016(1). https://doi.org/10.1186/s13640-016-0145-2

Yao, L., Liu, Y., & Huang, S. (2016). Spatio-temporal information for human action recognition. EURASIP Journal on Image and Video Processing, 2016(1). https://doi.org/10.1186/s13640-016-0145-2

Zhang, B., Yang, Y., Chen, C., Yang, L., Han, J., & Shao, L. (2017). Action recognition using 3D histograms of texture and a multi-class boosting classifier. Paper presented at the IEEE Transactions on Image Processing, 26(10), 4648–4660. https://doi.org/10.1109/ tip.2017.2718189

Zhang, B., Yang, Y., Chen, C., Yang, L., Han, J., & Shao, L. (2017). Action recognition using 3D histograms of texture and a multi-class boosting classifier. Paper presented at the IEEE Transactions on Image Processing, 26(10), 4648–4660. https://doi.org/10.1109/ tip.2017.2718189

Zhen, X., Zheng, F., Shao, L., Cao, X., & Xu, D. (2017). Supervised local descriptor learning for human action recognition. Paper presented at the IEEE Transactions on Multimedia, 19(9), 2056–2065. https://doi.org/10.1109/tmm.2017.2700204

Zhen, X., Zheng, F., Shao, L., Cao, X., & Xu, D. (2017). Supervised local descriptor learning for human action recognition. Paper presented at the IEEE Transactions on Multimedia, 19(9), 2056–2065. https://doi.org/10.1109/tmm.2017.2700204

Zhu, S., & Xia, L. (2015). Human action recognition based on fusion features extraction of adaptive background subtraction and optical flow model. Mathematical Problems in Engineering, 1–11. https://doi.org/10.1155/2015/387464

Downloads

Published

31-03-2019

How to Cite

Basavaiah, J., & Patil, C. M. (2019). Human Activity Detection and Action Recognition in Videos Using Convolutional Neural Networks. Journal of Information and Communication Technology, 19(2), 157-183. https://doi.org/10.32890/jict2020.19.2.6

Research impact

Harvested 2026-09-25
0 citations recorded so far

Counts differ between services because each indexes a different body of literature. None of them is the whole picture.

Identifiers DOI 10.32890/jict2020.19.2.6

Most read articles by the same author(s)