DocumentCode
153406
Title
Empirical Evaluation of CRF-Based Bibliography Extraction from Reference Strings
Author
Ohta, Masaya ; Arauchi, Daiki ; Takasu, Atsuhiro ; Adachi, Jun
Author_Institution
Okayama Univ., Okayama, Japan
fYear
2014
fDate
7-10 April 2014
Firstpage
287
Lastpage
292
Abstract
This paper reports an empirical evaluation of a CRF-based bibliography parser we have developed for reference strings of research papers. The parser uses a conditional random field (CRF) to estimate the correct bibliographic label such as an author´s name and a title for each token in a reference string. We applied the parser specifically designed for reference strings to three academic journals, an English one and two Japanese ones, published in Japan. Experiments showed (i) the parser correctly parsed from 90% to 94% of reference strings depending on the kinds of journals used and (ii) segmentation errors induced by tokenization considerably degraded the final parsing accuracies. This paper also discusses some future directions of the bibliography extraction based on a detailed analysis of the experiments.
Keywords
bibliographic systems; citation analysis; grammars; statistical analysis; string matching; CRF-based bibliography extraction; CRF-based bibliography parser; English academic journals; Japanese academic journals; author name; conditional random field; correct bibliographic label estimation; empirical evaluation; reference strings; research papers; segmentation errors; token title; tokenization; Accuracy; Bibliographies; Dictionaries; Feature extraction; Hidden Markov models; Labeling; Tagging; bibliography extraction; citation parsing; conditional random field; evaluation; metadata;
fLanguage
English
Publisher
ieee
Conference_Titel
Document Analysis Systems (DAS), 2014 11th IAPR International Workshop on
Conference_Location
Tours
Print_ISBN
978-1-4799-3243-6
Type
conf
DOI
10.1109/DAS.2014.64
Filename
6831015
Link To Document