• DocumentCode
    2988256
  • Title

    Drosophila GRAIL: an intelligent system for gene recognition in Drosophila DNA sequences

  • Author

    Xu, Ying ; Helt, Gregg ; Einstein, J. Ralph ; Rubin, Gerry ; Uberbacher, Edward C.

  • Author_Institution
    Div. of Comput. Sci. & Math., Oak Ridge Nat. Lab., TN, USA
  • fYear
    1995
  • fDate
    29-31 May 1995
  • Firstpage
    128
  • Lastpage
    135
  • Abstract
    An AI-based system for gene recognition in Drosophila DNA sequences was designed and implemented. The system consists of two main modules, one for coding exon recognition and one for single gene model construction. The exon recognition module finds a coding exon by recognition of its splice junctions (or translation start) and coding potential. The core of this module is a set of neural networks which evaluate an exon candidate for the possibility of being a true coding exon using the “recognized” splice junction (or translation start) and coding signals. The recognition process consists of four steps: generation of an exon candidate pool, elimination of improbable candidates using heuristic rules, candidate evaluation by trained neural networks, and candidate cluster resolution and final exon prediction. The gene model construction module takes as input the clustered exon candidates and builds a “best” possible single gene model using an efficient dynamic programming algorithm. 129 Drosophila sequences consisting of 441 coding exons including 216358 coding bases were extracted from GenBank and used to build statistical matrices and to train the neural networks. On this training set the system recognized 97% of the coding messages and predicted only 5% false messages. Among the “correctly” predicted exons, 68% match the actual exon exactly and 96% have at least one edge predicted correctly. On an independent test set consisting of 30 Drosophila sequences, the system recognized 96% of the coding messages and predicted 7% false messages
  • Keywords
    DNA; biology computing; genetics; knowledge based systems; neural nets; pattern recognition; sequences; Drosophila DNA sequences; Drosophila GRAIL; GenBank; candidate evaluation; cluster resolution; coding exon; dynamic programming; exon recognition; gene recognition; intelligent system; neural networks; single gene model construction; splice junctions; statistical matrices; Clustering algorithms; DNA; Dynamic programming; Heuristic algorithms; Intelligent systems; Modular construction; Neural networks; Sequences; Signal resolution; System testing;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Intelligence in Neural and Biological Systems, 1995. INBS'95, Proceedings., First International Symposium on
  • Conference_Location
    Herndon, VA
  • Print_ISBN
    0-8186-7116-5
  • Type

    conf

  • DOI
    10.1109/INBS.1995.404269
  • Filename
    404269