DocumentCode
2859002
Title
Tracking Humans using Multi-modal Fusion
Author
Zou, Xiaotao ; Bhanu, Bir
Author_Institution
University of California, Riverside
fYear
2005
fDate
25-25 June 2005
Firstpage
4
Lastpage
4
Abstract
Human motion detection plays an important role in automated surveillance systems. However, it is challenging to detect non-rigid moving objects (e.g. human) robustly in a cluttered environment. In this paper, we compare two approaches for detecting walking humans using multi-modal measurements- video and audio sequences. The first approach is based on the Time-Delay Neural Network (TDNN), which fuses the audio and visual data at the feature level to detect the walking human. The second approach employs the Bayesian Network (BN) for jointly modeling the video and audio signals. Parameter estimation of the graphical models is executed using the Expectation-Maximization (EM) algorithm. And the location of the target is tracked by the Bayes inference. Experiments are performed in several indoor and outdoor scenarios: in the lab, more than one person walking, occlusion by bushes etc. The comparison of performance and efficiency of the two approaches are also presented.
Keywords
Anthropometry; Computer vision; Fuses; Humans; Legged locomotion; Motion detection; Neural networks; Object detection; Robustness; Surveillance;
fLanguage
English
Publisher
ieee
Conference_Titel
Computer Vision and Pattern Recognition - Workshops, 2005. CVPR Workshops. IEEE Computer Society Conference on
Conference_Location
San Diego, CA, USA
ISSN
1063-6919
Print_ISBN
0-7695-2372-2
Type
conf
DOI
10.1109/CVPR.2005.545
Filename
1565299
Link To Document