A Compression Model for DNA Multiple Sequence Alignment Blocks

Matos Luis M O<sup>*</sup>; Pratas Diogo; Pinho Armando J

doi:10.1109/TIT.2012.2236605

摘要

A particularly voluminous dataset in molecular genomics, known as whole genome alignments, has gained considerable importance over the last years. In this paper, we propose a compression modeling approach for the multiple sequence alignment (MSA) blocks, which make up most of these datasets. Our method is based on a mixture of finite-context models. Contrarily to other recent approaches, it addresses both the DNA bases and gap symbols at once, better exploring the existing correlations. For comparison with previous methods, our algorithm was tested in the multiz28way dataset. On average, it attained 0.94 bits per symbol, approximately 7% better than the previous best, for a similar computational complexity. We also tested the model in the most recent dataset, multiz46way. In this dataset, that contains alignments of 46 different species, our compression model achieved an average of 0.72 bits per MSA block symbol.

出版日期2013-5

全文

访问全文

收藏分享被引浏览

更新时间：2017-04-25 10:35

A Compression Model for DNA Multiple Sequence Alignment Blocks

摘要

全文

产品服务

站内浏览

服务支持

联系方式

科研之友