• DocumentCode
    8202
  • Title

    Representing and Retrieving Video Shots in Human-Centric Brain Imaging Space

  • Author

    Junwei Han ; Xiang Ji ; Xintao Hu ; Dajiang Zhu ; Kaiming Li ; Xi Jiang ; Guangbin Cui ; Lei Guo ; Tianming Liu

  • Author_Institution
    Sch. of Autom., Northwestern Polytech. Univ., Xi´an, China
  • Volume
    22
  • Issue
    7
  • fYear
    2013
  • fDate
    Jul-13
  • Firstpage
    2723
  • Lastpage
    2736
  • Abstract
    Meaningful representation and effective retrieval of video shots in a large-scale database has been a profound challenge for the image/video processing and computer vision communities. A great deal of effort has been devoted to the extraction of low-level visual features, such as color, shape, texture, and motion for characterizing and retrieving video shots. However, the accuracy of these feature descriptors is still far from satisfaction due to the well-known semantic gap. In order to alleviate the problem, this paper investigates a novel methodology of representing and retrieving video shots using human-centric high-level features derived in brain imaging space (BIS) where brain responses to natural stimulus of video watching can be explored and interpreted. At first, our recently developed dense individualized and common connectivity-based cortical landmarks (DICCCOL) system is employed to locate large-scale functional brain networks and their regions of interests (ROIs) that are involved in the comprehension of video stimulus. Then, functional connectivities between various functional ROI pairs are utilized as BIS features to characterize the brain´s comprehension of video semantics. Then an effective feature selection procedure is applied to learn the most relevant features while removing redundancy, which results in the formation of the final BIS features. Afterwards, a mapping from low-level visual features to high-level semantic features in the BIS is built via the Gaussian process regression (GPR) algorithm, and a manifold structure is then inferred, in which video key frames are represented by the mapped feature vectors in the BIS. Finally, the manifold-ranking algorithm concerning the relationship among all data is applied to measure the similarity between key frames of video shots. Experimental results on the TRECVID 2005 dataset demonstrate the superiority of the proposed work in comparison with traditional methods.
  • Keywords
    biomedical MRI; brain; cognition; neurophysiology; video retrieval; DICCCOL; Gaussian process regression algorithm; TRECVID 2005 dataset; common connectivitybased cortical landmarks system; functional connectivities; high-level semantic features; human centric brain imaging space; human centric high-level features; large scale database; large-scale functional brain networks; meaningful representation; natural stimulus; regions of interests; retrieving video shot; video shot effective retrieval; video stimulus comprehension; video watching; Brain imaging space; Gaussian process regression; functional magnetic resonance imaging; video shot retrieval; Algorithms; Brain Mapping; Humans; Image Processing, Computer-Assisted; Magnetic Resonance Imaging; Regression Analysis; Semantics; Video Recording; Young Adult;
  • fLanguage
    English
  • Journal_Title
    Image Processing, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1057-7149
  • Type

    jour

  • DOI
    10.1109/TIP.2013.2256919
  • Filename
    6494293