Segmenting DNA sequence into `words'
2012-02-12Unverified0· sign in to hype
Wang Liang
Unverified — Be the first to reproduce this paper.
ReproduceAbstract
This paper presents a novel method to segment/decode DNA sequences based on n-grams statistical language model. Firstly, we find the length of most DNA 'words' is 12 to 15 bps by analyzing the genomes of 12 model species. Then we design an unsupervised probability based approach to segment the DNA sequences. The benchmark of segmenting method is also proposed.