• DocumentCode
    1376493
  • Title

    Photo-realistic talking-heads from image samples

  • Author

    Cosatto, Eric ; Graf, Hans Peter

  • Author_Institution
    AT&T Bell Labs.-Res., Red Bank, NJ, USA
  • Volume
    2
  • Issue
    3
  • fYear
    2000
  • fDate
    9/1/2000 12:00:00 AM
  • Firstpage
    152
  • Lastpage
    163
  • Abstract
    This paper describes a system for creating a photo-realistic model of the human head that can be animated and lip-synched from phonetic transcripts of text. Combined with a state-of-the-art text-to-speech synthesizer (TTS), it generates video animations of talking heads that closely resemble real people. To obtain a naturally looking head, we choose a “data-driven” approach. We record a talking person and apply image recognition to extract automatically bitmaps of facial parts. These bitmaps are normalized and parameterized before being entered into a database. For synthesis, the TTS provides the audio track, as well as the phonetic transcript from which trajectories in the space of parameterized bitmaps are computed for all facial parts. Sampling these trajectories and retrieving the corresponding bitmaps from the database produces animated facial parts. These facial parts are then projected and blended onto an image of the whole head using its pose information. This talking head model can produce new never recorded speech of the person who was originally recorded. Talking-head animations of this type are useful as a front-end for agents and avatars in multimedia applications such as virtual operators, virtual announcers, help desks, educational, and expert systems
  • Keywords
    computer animation; face recognition; multimedia computing; realistic images; software agents; speech synthesis; agents; audio track; avatars; computer vision; data-driven approach; educational systems; expert systems; face recognition; facial animation; facial parts; help desks; human head; image recognition; image samples; multimedia applications; phonetic transcripts; photo-realistic talking-heads; pose information; sample-based image synthesis; talking-head animations; text-to-speech synthesizer; video animations; virtual announcers; virtual operators; Animation; Humans; Image databases; Image recognition; Image sampling; Information retrieval; Magnetic heads; Speech synthesis; Synthesizers; Trajectory;
  • fLanguage
    English
  • Journal_Title
    Multimedia, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1520-9210
  • Type

    jour

  • DOI
    10.1109/6046.865480
  • Filename
    865480