Title :
Geotagging with local lexicons to build indexes for textually-specified spatial data
Author :
Lieberman, Michael D. ; Samet, Hanan ; Sankaranarayanan, Jagan
Author_Institution :
Dept. of Comput. Sci., Univ. of Maryland, College Park, MD, USA
Abstract :
The successful execution of location-based and feature-based queries on spatial databases requires the construction of spatial indexes on the spatial attributes. This is not simple when the data is unstructured as is the case when the data is a collection of documents such as news articles, which is the domain of discourse, where the spatial attribute consists of text that can be (but is not required to be) interpreted as the names of locations. In other words, spatial data is specified using text (known as a toponym) instead of geometry, which means that there is some ambiguity involved. The process of identifying and disambiguating references to geographic locations is known as geotagging and involves using a combination of internal document structure and external knowledge, including a document-independent model of the audience´s vocabulary of geographic locations, termed its spatial lexicon. In contrast to previous work, a new spatial lexicon model is presented that distinguishes between a global lexicon of locations known to all audiences, and an audience-specific local lexicon. Generic methods for inferring audiences´ local lexicons are described. Evaluations of this inference method and the overall geotagging procedure indicate that establishing local lexicons cannot be overlooked, especially given the increasing prevalence of highly local data sources on the Internet, and will enable the construction of more accurate spatial indexes.
Keywords :
Internet; data mining; geographic information systems; visual databases; Internet; document-independent model; external knowledge; feature-based queries; generic methods; geographic locations; geotagging; inference method; internal document structure; local lexicons; location-based queries; spatial databases; spatial indexes; spatial lexicon model; textually-specified spatial data; Automation; Computer science; Educational institutions; Frequency; Geometry; Influenza; Internet; Spatial databases; Spatial indexes; Vocabulary;
Conference_Titel :
Data Engineering (ICDE), 2010 IEEE 26th International Conference on
Conference_Location :
Long Beach, CA
Print_ISBN :
978-1-4244-5445-7
Electronic_ISBN :
978-1-4244-5444-0
DOI :
10.1109/ICDE.2010.5447903