DocumentCode :
3462114
Title :
Audio-visual synchrony for detection of monologues in video archives
Author :
Iyengar, G. ; Nock, H.J. ; Neti, C.
Author_Institution :
IBM Thomas J. Watson Res. Center, Yorktown Heights, NY, USA
Volume :
5
fYear :
2003
fDate :
6-10 April 2003
Abstract :
We present our approach to detecting monologues in video shots. A monologue shot is defined as a shot containing a talking person in the video channel with the corresponding speech in the audio channel. Whilst motivated by the TREC 2002 Video Retrieval Track (VT02), the underlying approach of synchrony between audio and video signals is also applicable for voice and face-based biometrics, assessing lip-synchronization quality in movie editing, and for speaker localization in video. Our approach is envisioned as a two part scheme. We first detect the occurrence of speech and face in a video shot. In shots containing both speech and a face, we distinguish monologue shots as those shots where the speech and facial movements are synchronized. To measure the synchrony between speech and facial movements we use a mutual-information based measure. Experiments with the VT02 corpus indicate that using synchrony, the average precision improves by more than 50% relative compared to using face and speech information alone. Our synchrony based monologue detector submission had the best average precision performance (in VT02) amongst 18 different submissions.
Keywords :
audio-visual systems; biometrics (access control); image recognition; image retrieval; speech recognition; synchronisation; video databases; video signal processing; audio-visual synchrony; face-based biometrics; lip-synchronization quality; monologue detection; speaker localization; video archives; video retrieval; voice-based biometrics; Biometrics; Data analysis; Event detection; Face detection; Gunshot detection systems; Motion pictures; Noise robustness; Random variables; Speech recognition; Training data;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03). 2003 IEEE International Conference on
ISSN :
1520-6149
Print_ISBN :
0-7803-7663-3
Type :
conf
DOI :
10.1109/ICASSP.2003.1200085
Filename :
1200085
Link To Document :
بازگشت