Now accepting submissions for the upcoming volume
Electronic and Acoustic Engineering
ISSN 2524-2725 · e‑ISSN 2617-0965 Open Access · CC BY-NC 4.0
Vol. 2 · Issue 3 · 2019 Jun 28, 2019 Electronic Systems and Signals

Action and movements recognition methods

DS
Denys Soldatov Corresponding National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute” kingit@bk.ru Ukraine
Pages27-33 PublishedJun 28, 2019 LicenseOpen Access
EAE 3 VOL 2 · 3
VOL 2 · NO 3 · 2019 View issue

Abstract

The article describes the formulation of the problem of recognition of the movements of objects in a video sequence, the stages of its solution, the analysis of the basic methods of each of the stages. A wide range of applications and growing requirements on the quality of recognition determines the relevance of the study. The process of action recognition and detection begins with extracting useful features, from the input video sequence. Features are then processed through a classier to identify the action class (for example, running, walking, jumping, various gestures). The article describes the main feature descriptors, in the filter-based category: histogram of oriented gradients, cuboid descriptor, scale-invariant feature transform, gradient location-orientation histogram, local trinary patterns, and spatiotemporal patches, optical flow-based descriptors: histograms of optical flow, the motion boundary histogram, dense trajectory, convolutional neural network-based descriptors. Some algorithms require the extraction of primitive features and further refinement of the auxiliary features before they can be passed to the classifier. Examples of the use of specialized primitive features are methods based on silhouettes / contours and methods based on object tracking. There are methods for classifying extracted features, including the following: support vector machines, adaptive boost, artificial neural networks, convolutional neural networks. The key difficulties arising in solving the problem are considered. There are ways to compare various methods. One of the ways to draw comparisons is to quantitatively evaluate each approach on the same database with the same protocol. From simple KTH datasets and Weizmannnd to Carnegie Mellon University Crowded Videos dataset and Microsoft Research Action Group dataset to more complex video conditions and large-scale UCF101 and ActivityNet datasets. Existing approaches to recognition of motion in video sequences are analyzed. The article reveals characteristics, strengths and weaknesses of the various methods of detecting features and their classification. Leading methods that show the best results widely use convolutional neural networks. One of such methods is a spatio-temporal graph convolutional neural network for action recognition based on the object's skeleton. A method for further research and improvement was chosen.

Keywords

References

  1. Y. Du, F. Chen, and W. Xu, “Human interaction representation and recognition through motion decomposition,” IEEE Signal Processing Letters, vol. 14, no. 12, pp. 952–955, 2007. DOI: 10.1109/LSP.2007.908035
  2. C. Schüldt, I. Laptev, and B. Caputo, “Recognizing human actions: A local SVM approach,” in Proceedings - International Conference on Pattern Recognition, 2004, vol. 3, pp. 32–36. DOI: 10.1109/ICPR.2004.1334462
  3. L. Gorelick, M. Blank, E. Shechtman, M. Irani, and R. Basri, “Actions as space-time shapes,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 29, no. 12, pp. 2247–2253, 2007. DOI: 10.1109/TPAMI.2007.70711
  4. Y. Ke, R. Sukthankar, and M. Hebert, “Event detection in crowded videos,” in Proceedings of the IEEE International Conference on Computer Vision, 2007, pp. 1–8. DOI: 10.1109/ICCV.2007.4409011
  5. J. Yuan, Z. Liu, and Y. Wu, “Discriminative subvolume search for efficient action detection,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, pp. 2442–2449. DOI: 10.1109/CVPRW.2009.5206671
  6. I. Laptev and P. Pérez, “Retrieving actions in movies,” in Proceedings of the IEEE International Conference on Computer Vision, 2007, pp. 1–8. DOI: 10.1109/ICCV.2007.4409105
  7. M. D. Rodriguez, J. Ahmed, and M. Shah, “Action MACH: A spatio-temporal maximum average correlation height filter for action recognition,” in 26th IEEE Computer Society Conference on Computer Vision and Pattern Recognition, CVPR, 2008. DOI: 10.1109/CVPR.2008.4587727
  8. I. Laptev, M. Marszałek, C. Schmid, and B. Rozenfeld, “Learning realistic human actions from movies,” in 26th IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2008. DOI: 10.1109/CVPR.2008.4587756
  9. M. Marszałek, I. Laptev, and C. Schmid, “Actions in context,” in IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2009, pp. 2929–2936. DOI: 10.1109/CVPRW.2009.5206557
  10. K. Soomro, A. R. Zamir, and M. Shah, “UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild,” 2012. URL: http://arxiv.org/abs/1212.0402
  11. C. Snoek, B. Ghanem, J.C. Niebles, F.C. Heilbron, W. Barrios, V. Escorcia, and P. Mettes. ActivityNet: A Large-Scale Activity Recognition Challenge.
  12. S. M. Kang and R. P. Wildes, “Review of Action Recognition and Detection Methods,” 2016. URL: http://arxiv.org/abs/1610.06906
  13. J. Wang, P. Liu, M. F. H. She, A. Kouzani, and S. Nahavandi, “Supervised learning probabilistic Latent Semantic Analysis for human motion analysis,” Neurocomputing, vol. 100, pp. 134–143, 2013. DOI: 10.1016/j.neucom.2011.10.033
  14. A. Klaeser, M. Marszalek, and C. Schmid, “A Spatio-Temporal Descriptor Based on 3D-Gradients,” in British Machine Vision Conference, 2012, p. 99.1-99.10. DOI: 10.5244/c.22.99
  15. P. Dollár, V. Rabaud, G. Cottrell, and S. Belongie, “Behavior recognition via sparse spatio-temporal features,” in Proceedings - 2nd Joint IEEE International Workshop on Visual Surveillance and Performance Evaluation of Tracking and Surveillance, VS-PETS, 2005, pp. 65–72. DOI: 10.1109/VSPETS.2005.1570899
  16. D. Lowe, “Distinctive Image Features from Scale-Invariant Keypoints,” Int. J. Comput. Vis., vol. 60, no. 2, pp. 91–110, 2004. DOI: 10.1023/B:VISI.0000029664.99615.94
  17. K. Mikolajczk and C. Schmid, “A performance of local descriptors,” IEEE Conf. Comput. Vis. Pattern Recognit., vol. 27, no. 10, pp. 1615–1630, 2003. DOI: 10.1109/TPAMI.2005.188
  18. L. Yeffet and L. Wolf, “Local trinary patterns for human action recognition,” in Proceedings of the IEEE International Conference on Computer Vision, 2009, pp. 492–497. DOI: 10.1109/ICCV.2009.5459201
  19. E. Shechtman and M. Irani, “Space-time behavior-based correlation - OR - How to tell if two underlying motion fields are similar without computing them?,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 29, no. 11, pp. 2045–2056, 2007. DOI: 10.1109/TPAMI.2007.1119
  20. H. Ning, T. X. Han, D. B. Walther, M. Liu, and T. S. Huang, “Hierarchical space-time model enabling efficient search for human actions,” IEEE Trans. Circuits Syst. Video Technol., vol. 19, no. 6, pp. 808–820, 2009. DOI: 10.1109/TCSVT.2009.2017399
  21. O. Chomat and J. L. Crowley, “Probabilistic recognition of activity using local appearance,” 2003, pp. 104–109. DOI: 10.1109/cvpr.1999.784616
  22. J. M. Gryn, R. P. Wildes, and J. K. Tsotsos, “Detecting motion patterns via direction maps with application to surveillance,” Comput. Vis. Image Underst., vol. 113, no. 2, pp. 291–307, 2009. DOI: 10.1016/j.cviu.2008.10.006
  23. A. A. Efros, A. C. Berg, G. Mori, and J. Malik, “Recognizing action at a distance,” in IEEE International Conference on Computer Vision, 2004, pp. 726–733. DOI: 10.1109/iccv.2003.1238420
  24. A. Fathi and G. Mori, “Action recognition by learning mid-level motion features,” in 26th IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2008. DOI: 10.1109/CVPR.2008.4587735
  25. N. Dalal, B. Triggs, and C. Schmid, “Human detection using oriented histograms of flow and appearance,” Lect. Notes Comput. Sci. (including Subser. Lect. Notes Artif. Intell. Lect. Notes Bioinformatics), vol. 3952 LNCS, pp. 428–441, 2006. DOI: 10.1007/11744047_33
  26. H. Wang, A. Kläser, C. Schmid, and C. L. Liu, “Action recognition by dense trajectories,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2011, pp. 3169–3176. DOI: 10.1109/CVPR.2011.5995407
  27. H. Wang and C. Schmid, “Action recognition with improved trajectories,” Proc. IEEE Int. Conf. Comput. Vis., pp. 3551–3558, 2013. DOI: 10.1109/ICCV.2013.441
  28. H. Bay, T. Tuytelaars, and L. Van Gool, “SURF: Speeded up robust features,” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2006, vol. 3951 LNCS, pp. 404–417. DOI: 10.1007/11744023_32
  29. J. Shi and C. Tomasi, “Good features to track,” in Proceedings of IEEE Conference on Computer Vision and Pattern Recognition CVPR-94, 1994, pp. 593–600. DOI: 10.1109/CVPR.1994.323794
  30. H. Bilen, B. Fernando, E. Gavves, A. Vedaldi, and S. Gould, “Dynamic Image Networks for Action Recognition,” in IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 3034–3042. DOI: 10.1109/cvpr.2016.331
  31. X. Wang, A. Farhadi, and A. Gupta, “Actions ~ Transformations,” in IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 2658–2667. URL: http://arxiv.org/abs/1512.00795
  32. Z. Shou, D. Wang, and S.-F. Chang, “Temporal Action Localization in Untrimmed Videos via Multi-stage CNNs,” in IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 1049–1058. URL: http://arxiv.org/abs/1601.02129
  33. J. Yue et al., “Beyond short snippets: Deep networks for video classification,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2015, vol. 07–12–June, pp. 4694–4702. DOI: 10.1109/CVPR.2015.7299101
  34. S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei, “End-to-end Learning of Action Detection from Frame Glimpses in Videos,” in IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 2678–2687. DOI: 10.1109/CVPR.2016.293
  35. A. Basharat, A. Gritai, and M. Shah, “Learning object motion patterns for anomaly detection and improved object detection,” in 26th IEEE Conference on Computer Vision and Pattern Recognition, CVPR, 2008, pp. 1–8. DOI: 10.1109/CVPR.2008.4587510
  36. C. Fanti, L. Zelnik-Manor, and P. Perona, “Hybrid models for human motion recognition,” in Proceedings - 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, CVPR 2005, 2005, vol. I, pp. 1166–1173. DOI: 10.1109/CVPR.2005.179
  37. S. Yeung, O. Russakovsky, G. Mori, and L. Fei-Fei, “End-to-end Learning of Action Detection from Frame Glimpses in Videos,” 2015. URL: http://arxiv.org/abs/1511.06984
  38. S. Yan, Y. Xiong, and D. Lin, “Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition,” Dep. Inf. Eng. Chinese Univ. Hong Kong, 2018. URL: http://arxiv.org/abs/1801.07455

License

CCBY-NC 4.0
Creative Commons Attribution 4.0 International

This work is openly licensed — share and adapt freely with attribution to the authors and the journal. View license terms ↗

§ 06 — Related

Similar articles in this journal

Related peer-reviewed studies published in this journal.
View all issues

Similar Articles

1-10 of 40

You may also start an advanced similarity search for this article.