Title :
Truncation of protein sequences for fast profile alignment with application to subcellular localization
Author :
Mak, Man-Wai ; Wang, Wei ; Kung, Sun-Yuan
Author_Institution :
Dept. of Electron. & Inf. Eng., Hong Kong Polytech. Univ., Hong Kong, China
Abstract :
We have recently found that the computation time of homology-based subcellular localization can be substantially reduced by aligning profiles up to the cleavage site positions of signal peptides, mitochondrial targeting peptides, and chloro-plast transit peptides [1]. While the method can reduce the profile alignment time by as much as 20 folds, it cannot reduce the computation time spent on creating the profiles. In this paper, we propose a new approach that can reduce both the profile creation time and profile alignment time. In the new approach, instead of cutting the profiles, we shorten the sequences by cutting them at the cleavage site locations. The shortened sequences are then presented to PSI-BLAST to compute the profiles. Experimental results and analysis of profile-alignment score matrices suggest that both profile creation time and profile alignment time can be reduced without sacrificing subcellular localization accuracy. Once a pairwise profile-alignment score matrix has been obtained, a one-vs-rest SVM classifier can be trained. To further reduce the training and recognition time of the classifier, we propose a perturbation discriminant analysis (PDA) technique. It was found that PDA enjoys a short training time as compared to the conventional SVM.
Keywords :
bioinformatics; molecular biophysics; molecular configurations; pattern classification; proteins; statistical analysis; support vector machines; PDA technique; PSI-BLAST; SVM classifier training; chloroplast transit peptides; fast profile alignment; homology based subcellular localization computation time; mitochondrial targeting peptides; peptide cleavage site positions; perturbation discriminant analysis; profile alignment score matrix analysis; profile alignment time reduction; profile creation time reduction; protein sequence truncation; signal peptides; Accuracy; Amino acids; Kernel; Personal digital assistants; Proteins; Support vector machines; Training; SVM; Subcellular localization; cleavage sites prediction; kernel discriminant analysis; profiles alignment; protein sequences;
Conference_Titel :
Bioinformatics and Biomedicine (BIBM), 2010 IEEE International Conference on
Conference_Location :
Hong Kong
Print_ISBN :
978-1-4244-8306-8
Electronic_ISBN :
978-1-4244-8307-5
DOI :
10.1109/BIBM.2010.5706548