Title :
Probabilistic Arithmetic Automata and Their Applications
Author :
Marschall, T. ; Herms, I. ; Kaltenbach, H. ; Rahmann, S.
Author_Institution :
Life Sci. Group, Centrum Wiskunde & Inf. (CWI), Amsterdam, Netherlands
Abstract :
We present a comprehensive review on probabilistic arithmetic automata (PAAs), a general model to describe chains of operations whose operands depend on chance, along with two algorithms to numerically compute the distribution of the results of such probabilistic calculations. PAAs provide a unifying framework to approach many problems arising in computational biology and elsewhere. We present five different applications, namely 1) pattern matching statistics on random texts, including the computation of the distribution of occurrence counts, waiting times, and clump sizes under hidden Markov background models; 2) exact analysis of window-based pattern matching algorithms; 3) sensitivity of filtration seeds used to detect candidate sequence alignments; 4) length and mass statistics of peptide fragments resulting from enzymatic cleavage reactions; and 5) read length statistics of 454 and IonTorrent sequencing reads. The diversity of these applications indicates the flexibility and unifying character of the presented framework. While the construction of a PAA depends on the particular application, we single out a frequently applicable construction method: We introduce deterministic arithmetic automata (DAAs) to model deterministic calculations on sequences, and demonstrate how to construct a PAA from a given DAA and a finite-memory random text model. This procedure is used for all five discussed applications and greatly simplifies the construction of PAAs. Implementations are available as part of the MoSDi package. Its application programming interface facilitates the rapid development of new applications based on the PAA framework.
Keywords :
arithmetic; biology computing; enzymes; hidden Markov models; pattern matching; probabilistic automata; IonTorrent sequencing reads; PAA; application programming interface; computational biology; deterministic arithmetic automata; enzymatic cleavage reactions; filtration seeds; finite-memory random text model; hidden Markov background models; pattern matching statistics; probabilistic arithmetic automata; Automata; Bioinformatics; Computational modeling; Hidden Markov models; Markov processes; Probabilistic logic; DNA sequencing; Probabilistic automaton; alignment seed; analysis of algorithms; clump; dynamic programming.; hidden Markov model; pattern matching; peptide mass fingerprinting; statistics; string algorithm; text model; Algorithms; Computational Biology; Markov Chains; Models, Biological; Models, Statistical; Pattern Recognition, Automated; Peptide Mapping; Sensitivity and Specificity; Sequence Analysis, DNA;
Journal_Title :
Computational Biology and Bioinformatics, IEEE/ACM Transactions on
DOI :
10.1109/TCBB.2012.109