Photo-realistic talking-heads from image samples

Author

Cosatto, Eric ; Graf, Hans Peter

Author_Institution

AT&T Bell Labs.-Res., Red Bank, NJ, USA

Volume

2

Issue

3

fYear

2000

fDate

9/1/2000 12:00:00 AM

Firstpage

152

Lastpage

163

Abstract

This paper describes a system for creating a photo-realistic model of the human head that can be animated and lip-synched from phonetic transcripts of text. Combined with a state-of-the-art text-to-speech synthesizer (TTS), it generates video animations of talking heads that closely resemble real people. To obtain a naturally looking head, we choose a “data-driven” approach. We record a talking person and apply image recognition to extract automatically bitmaps of facial parts. These bitmaps are normalized and parameterized before being entered into a database. For synthesis, the TTS provides the audio track, as well as the phonetic transcript from which trajectories in the space of parameterized bitmaps are computed for all facial parts. Sampling these trajectories and retrieving the corresponding bitmaps from the database produces animated facial parts. These facial parts are then projected and blended onto an image of the whole head using its pose information. This talking head model can produce new never recorded speech of the person who was originally recorded. Talking-head animations of this type are useful as a front-end for agents and avatars in multimedia applications such as virtual operators, virtual announcers, help desks, educational, and expert systems

Keywords

computer animation; face recognition; multimedia computing; realistic images; software agents; speech synthesis; agents; audio track; avatars; computer vision; data-driven approach; educational systems; expert systems; face recognition; facial animation; facial parts; help desks; human head; image recognition; image samples; multimedia applications; phonetic transcripts; photo-realistic talking-heads; pose information; sample-based image synthesis; talking-head animations; text-to-speech synthesizer; video animations; virtual announcers; virtual operators; Animation; Humans; Image databases; Image recognition; Image sampling; Information retrieval; Magnetic heads; Speech synthesis; Synthesizers; Trajectory;

fLanguage

English

Journal_Title

Multimedia, IEEE Transactions on

Publisher

ieee

ISSN

1520-9210

Type

jour

DOI

10.1109/6046.865480

Filename

865480