• DocumentCode
    3639683
  • Title

    Tools and methodologies for annotating syntax and named entities in the National Corpus of Polish

  • Author

    Jakub Waszczuk;Katarzyna Glowińska;Agata Savary;Adam Przepiórkowski

  • Author_Institution
    Institute of Computer Science, Polish Academy of Sciences, ul. Ordona 21, 01-237 Warsaw, Poland
  • fYear
    2010
  • Firstpage
    531
  • Lastpage
    539
  • Abstract
    The on-going project aiming at the creation of the National Corpus of Polish assumes several levels of linguistic annotation. We present the technical environment and methodological background developed for the three upper annotation levels: the level of syntactic words and groups, and the level of named entities. We show how knowledge-based platforms Spejd and Sprout are used for the automatic pre-annotation of the corpus, and we discuss some particular problems faced during the elaboration of the syntactic grammar, which contains over 800 rules and is one of the largest chunking grammars for Polish. We also show how the tree editor TrEd has been customized for manual post-editing of annotations, and for further revision of discrepancies. Our XML format converters and customized archiving repository ensure the automatic data flow and efficient corpus file management. We believe that this environment or substantial parts of it can be reused in or adapted for other corpus annotation tasks.
  • Keywords
    "Syntactics","Grammar","Manuals","Context","Semantics","Pragmatics","Computer science"
  • Publisher
    ieee
  • Conference_Titel
    Computer Science and Information Technology (IMCSIT), Proceedings of the 2010 International Multiconference on
  • ISSN
    2157-5525
  • Print_ISBN
    978-1-4244-6432-6
  • Electronic_ISBN
    2157-5533
  • Type

    conf

  • DOI
    10.1109/IMCSIT.2010.5680057
  • Filename
    5680057