DocumentCode
3424191
Title
An audio-visual fusion framework with joint dimensionality reducton
Author
Liu, Ming ; Fu, Yun ; Huang, Thomas S.
Author_Institution
Beckman Inst. for Adv. Sci. & Technol., Illinois Univ., Urbana, IL
fYear
2008
fDate
March 31 2008-April 4 2008
Firstpage
4437
Lastpage
4440
Abstract
By combining audio and visual modalities, the speech recognition systems achieve higher performance and robustness. The fusion strategies to this point are mainly three types: feature level fusion, model level fusion, and decision level fusion. In this paper, we present a novel audio-visual fusion framework, in which a joint dimensionality reduction approach is used to project the audio and visual features into more compact subspaces. With correlation preserving criteria, the representations of projected audio and visual features will be able to preserve the correlation conveyed in the original audio and visual feature space. At the same time, the better model efficiency is achieved in the more compact feature spaces. The experiments on audio-visual person verification demonstrate the efficiency and effectiveness of the proposed fusion framework.
Keywords
speaker recognition; video signal processing; audio-visual fusion framework; audio-visual person verification; decision level fusion; feature level fusion; joint dimensionality reduction; model level fusion; speech recognition; Acoustic noise; Automatic speech recognition; Availability; Feature extraction; Humans; Loudspeakers; Pattern recognition; Robustness; Speech recognition; Streaming media; Audio-visual fusion; audio-visual person verification; canonical correlation analysis; dimensionality reduction;
fLanguage
English
Publisher
ieee
Conference_Titel
Acoustics, Speech and Signal Processing, 2008. ICASSP 2008. IEEE International Conference on
Conference_Location
Las Vegas, NV
ISSN
1520-6149
Print_ISBN
978-1-4244-1483-3
Electronic_ISBN
1520-6149
Type
conf
DOI
10.1109/ICASSP.2008.4518640
Filename
4518640
Link To Document