• DocumentCode
    177933
  • Title

    Scale Coding Bag-of-Words for Action Recognition

  • Author

    Shahbaz Khan, F. ; Van De Weijer, J. ; Bagdanov, A.D. ; Felsberg, M.

  • Author_Institution
    Comput. Vision Lab., Linkoping Univ., Linkoping, Sweden
  • fYear
    2014
  • fDate
    24-28 Aug. 2014
  • Firstpage
    1514
  • Lastpage
    1519
  • Abstract
    Recognizing human actions in still images is a challenging problem in computer vision due to significant amount of scale, illumination and pose variation. Given the bounding box of a person both at training and test time, the task is to classify the action associated with each bounding box in an image. Most state-of-the-art methods use the bag-of-words paradigm for action recognition. The bag-of-words framework employing a dense multi-scale grid sampling strategy is the de facto standard for feature detection. This results in a scale invariant image representation where all the features at multiple-scales are binned in a single histogram. We argue that such a scale invariant strategy is sub-optimal since it ignores the multi-scale information available with each bounding box of a person. This paper investigates alternative approaches to scale coding for action recognition in still images. We encode multi-scale information explicitly in three different histograms for small, medium and large scale visual-words. Our first approach exploits multi-scale information with respect to the image size. In our second approach, we encode multi-scale information relative to the size of the bounding box of a person instance. In each approach, the multi-scale histograms are then concatenated into a single representation for action classification. We validate our approaches on the Willow dataset which contains seven action categories: interacting with computer, photography, playing music, riding bike, riding horse, running and walking. Our results clearly suggest that the proposed scale coding approaches outperform the conventional scale invariant technique. Moreover, we show that our approach obtains promising results compared to more complex state-of-the-art methods.
  • Keywords
    computer vision; feature selection; gesture recognition; image classification; image representation; image sampling; Willow dataset; bounding box; computer vision; dense multiscale grid sampling strategy; feature detection; human action recognition; multiscale information; scale coding bag-of-words; scale invariant image representation; single histogram; Encoding; Feature extraction; Histograms; Image coding; Image recognition; Image representation; Visualization;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Pattern Recognition (ICPR), 2014 22nd International Conference on
  • Conference_Location
    Stockholm
  • ISSN
    1051-4651
  • Type

    conf

  • DOI
    10.1109/ICPR.2014.269
  • Filename
    6976979