DocumentCode
2426630
Title
Meta Data Extraction from Linguistic Meeting Transcripts for the Annodex File Format
Author
Schremmer, Claudia ; Pfeiffer, Silvia
Author_Institution
CSIRO-ICT Centre
fYear
2005
fDate
12-14 Jan. 2005
Firstpage
405
Lastpage
412
Abstract
Semantic interpretation of the data distributed over the Internet is subject to major current research activity. The Continuous Media Web (CMWeb) extends the World Wide Web to time-continuously sampled data such as audio and video in regard to the searching, linking, and browsing functionality. The CMWeb technology is based the file format Annodex which streams the media content interspersed with markup in the Continuous Media Markup Language (CMML) format that contains information relevant to the whole media file, e.g., title, author, language as well as time-sensitive information, e.g., topics, speakers, time-sensitive hyperlinks. The CMML markup may be generated manually or automatically. This paper investigates the automatic extraction of meta data and markup information from complex linguistic annotations, which are annotated recordings collected for use in linguistic research. We are particularly interested in annotated recordings of meetings and teleconferences and see automatically generated CMML files and their corresponding Annodex streams as one way of viewing such recordings. The paper presents some experiments with generating Annodex files from hand-annotated meeting recordings.
Keywords
Annodex; Continuous Media Markup Language; Continuous Media Web; Linguistic Transcriptions; Markup; Meta Data; Data mining; Information retrieval; Internet; Markup languages; Natural languages; Search engines; Speech; Streaming media; Teleconferencing; Web sites;
fLanguage
English
Publisher
ieee
Conference_Titel
Multimedia Modelling Conference, 2005. MMM 2005. Proceedings of the 11th International
ISSN
1550-5502
Print_ISBN
0-7695-2164-9
Type
conf
DOI
10.1109/MMMC.2005.53
Filename
1386022
Link To Document