DocumentCode :
1376493
Title :
Photo-realistic talking-heads from image samples
Author :
Cosatto, Eric ; Graf, Hans Peter
Author_Institution :
AT&T Bell Labs.-Res., Red Bank, NJ, USA
Volume :
2
Issue :
3
fYear :
2000
fDate :
9/1/2000 12:00:00 AM
Firstpage :
152
Lastpage :
163
Abstract :
This paper describes a system for creating a photo-realistic model of the human head that can be animated and lip-synched from phonetic transcripts of text. Combined with a state-of-the-art text-to-speech synthesizer (TTS), it generates video animations of talking heads that closely resemble real people. To obtain a naturally looking head, we choose a “data-driven” approach. We record a talking person and apply image recognition to extract automatically bitmaps of facial parts. These bitmaps are normalized and parameterized before being entered into a database. For synthesis, the TTS provides the audio track, as well as the phonetic transcript from which trajectories in the space of parameterized bitmaps are computed for all facial parts. Sampling these trajectories and retrieving the corresponding bitmaps from the database produces animated facial parts. These facial parts are then projected and blended onto an image of the whole head using its pose information. This talking head model can produce new never recorded speech of the person who was originally recorded. Talking-head animations of this type are useful as a front-end for agents and avatars in multimedia applications such as virtual operators, virtual announcers, help desks, educational, and expert systems
Keywords :
computer animation; face recognition; multimedia computing; realistic images; software agents; speech synthesis; agents; audio track; avatars; computer vision; data-driven approach; educational systems; expert systems; face recognition; facial animation; facial parts; help desks; human head; image recognition; image samples; multimedia applications; phonetic transcripts; photo-realistic talking-heads; pose information; sample-based image synthesis; talking-head animations; text-to-speech synthesizer; video animations; virtual announcers; virtual operators; Animation; Humans; Image databases; Image recognition; Image sampling; Information retrieval; Magnetic heads; Speech synthesis; Synthesizers; Trajectory;
fLanguage :
English
Journal_Title :
Multimedia, IEEE Transactions on
Publisher :
ieee
ISSN :
1520-9210
Type :
jour
DOI :
10.1109/6046.865480
Filename :
865480
Link To Document :
بازگشت