Title :
BibPro: A Citation Parser Based on Sequence Alignment
Author :
Chen, Chien-Chin ; Yang, Kai-Hsiang ; Chen, Chuen-Liang ; Ho, Jan-Ming
Author_Institution :
Dept. of Comput. Sci. & Inf. Eng., Nat. Taiwan Univ., Taipei, Taiwan
Abstract :
Dramatic increase in the number of academic publications has led to growing demand for efficient organization of the resources to meet researchers´ needs. As a result, a number of network services have compiled databases from the public resources scattered over the Internet. However, publications by different conferences and journals adopt different citation styles. It is an interesting problem to accurately extract metadata from a citation string which is formatted in one of thousands of different styles. It has attracted a great deal of attention in research in recent years. In this paper, based on the notion of sequence alignment, we present a citation parser called BibPro that extracts components of a citation string. To demonstrate the efficacy of BibPro, we conducted experiments on three benchmark data sets. The results show that BibPro achieved over 90 percent accuracy on each benchmark. Even with citations and associated metadata retrieved from the web as training data, our experiments show that BibPro still achieves a reasonable performance.
Keywords :
Internet; citation analysis; meta data; BibPro; Internet; World Wide Web; academic publications; citation parser; citation string; metadata extraction; public resources; sequence alignment; Benchmark testing; Decision support systems; Information retrieval; Internet; Metadata; Resource management; Sequential analysis; Web and internet services; Data integration; digital libraries; information extraction; sequence alignment.;
Journal_Title :
Knowledge and Data Engineering, IEEE Transactions on
DOI :
10.1109/TKDE.2010.231