Segmenting DNA sequence into `words'

2012-02-12Unverified0· sign in to hype

Wang Liang

Unverified — Be the first to reproduce this paper.

Abstract

This paper presents a novel method to segment/decode DNA sequences based on n-grams statistical language model. Firstly, we find the length of most DNA 'words' is 12 to 15 bps by analyzing the genomes of 12 model species. Then we design an unsupervised probability based approach to segment the DNA sequences. The benchmark of segmenting method is also proposed.

Tasks

Language Modeling Language Modelling

Segmenting DNA sequence into `words'

Abstract

Tasks

Reproductions