DocumentCode
1791576
Title
Topic similarity networks: Visual analytics for large document sets
Author
Maiya, Arun S. ; Rolfe, Robert M.
Author_Institution
Inst. for Defense Anal., Alexandria, VA, USA
fYear
2014
fDate
27-30 Oct. 2014
Firstpage
364
Lastpage
372
Abstract
We investigate ways in which to improve the interpretability of LDA topic models by better analyzing and visualizing their outputs. We focus on examining what we refer to as topic similarity networks: graphs in which nodes represent latent topics in text collections and links represent similarity among topics. We describe efficient and effective approaches to both building and labeling such networks. Visualizations of topic models based on these networks are shown to be a powerful means of exploring, characterizing, and summarizing large collections of unstructured text documents. They help to “tease out” non-obvious connections among different sets of documents and provide insights into how topics form larger themes. We demonstrate the efficacy and practicality of these approaches through two case studies: 1) NSF grants for basic research spanning a 14 year period and 2) the entire English portion of Wikipedia.
Keywords
data analysis; data visualisation; natural language processing; NSF; Wikipedia; large document sets; topic similarity network; unstructured text documents; visual analytics; Big data; Communities; Data visualization; Labeling; Probability distribution; Proteins; Visualization;
fLanguage
English
Publisher
ieee
Conference_Titel
Big Data (Big Data), 2014 IEEE International Conference on
Conference_Location
Washington, DC
Type
conf
DOI
10.1109/BigData.2014.7004253
Filename
7004253
Link To Document