DocumentCode
3659545
Title
Novel self-learning based crawling and data mining for automatic information extraction
Author
Arun Kumar A V;Hemant Kumar Rath;Shameemraj M Nadaf;Anantha Simha
Author_Institution
CTO Networks Lab, Tata Consultancy Services, Bangalore, India
fYear
2015
Firstpage
732
Lastpage
738
Abstract
In this paper, we propose techniques using a novel combination of self-learning based crawling and rule based data mining. Using the crawling techniques smaller relevant data sets can be obtained pertaining to a domain from multi-dimensional data sets available in on-line as well as off-line sources. We then process the crawled data sets and mine to extract meaningful information. Our techniques are generic in nature and can be used for automatic information extraction in different domains such as biomedical, health-care, enterprise infrastructure planning, etc. The proposed schemes are of reduced time, space and processor complexity due to the assisted and learning nature of the crawling. The data mining is based on configurable classification rules and decision trees, which are scalable and easy to implement in practice. We evaluate our proposed techniques through Java based implementation and integration with TCS in-house enterprise network design tool NetDes.
Keywords
"Data mining","Crawlers","Complexity theory","Vegetation","Java","Planning"
Publisher
ieee
Conference_Titel
Advances in Computing, Communications and Informatics (ICACCI), 2015 International Conference on
Print_ISBN
978-1-4799-8790-0
Type
conf
DOI
10.1109/ICACCI.2015.7275698
Filename
7275698
Link To Document