A Parallel Algorithm for the Fixed-length Approximate String Matching Problem for High Throughput Sequencing Technologies

Costas S. Iliopoulos; Laurent Mouchard; Solon P. Pissis

doi:10.3233/978-1-60750-530-3-150

A Parallel Algorithm for the Fixed-length Approximate String Matching Problem for High Throughput Sequencing Technologies

Costas S. Iliopoulos, Laurent Mouchard, Solon P. Pissis

Informatics

Research output: Chapter in Book/Report/Conference proceeding › Conference paper

2 Citations (Scopus)

Abstract

The approximate string matching problem is to find all locations at which a query of length m matches a substring of a text of length n with k-or-fewer differences. Nowadays, with the advent of novel high throughput sequencing technologies, the approximate string matching algorithms are used to identify similarities, molecular functions and abnormalities in DNA sequences. We consider a generalization of this problem, the fixed-length approximate string matching problem: given a text t, a pattern ρ and an integer ℓ, compute the optimal alignment of all substrings of ρ of length ℓ and a substring of t. We present a practical parallel algorithm of comparable simplicity that requires only time, where w is the word size of the machine (e.g. 32 or 64 in practice) and p the number of processors, by virtue of computing a bit representation of the relocatable dynamic programming matrix for the problem. Thus the algorithm's performance is independent of k and the alphabet size |Σ|.

Original language	Undefined/Unknown
Title of host publication	Proceedings of the International Conference on Parallel Computing (PARCO 2009)
Publisher	IOS Press
Pages	150-157
Number of pages	8
Volume	19
DOIs	https://doi.org/10.3233/978-1-60750-530-3-150
Publication status	Published - 2010

Access to Document

10.3233/978-1-60750-530-3-150

Cite this

@inbook{42661f7ad3c8415e8e2094e4c9317188,

title = "A Parallel Algorithm for the Fixed-length Approximate String Matching Problem for High Throughput Sequencing Technologies",

abstract = "The approximate string matching problem is to find all locations at which a query of length m matches a substring of a text of length n with k-or-fewer differences. Nowadays, with the advent of novel high throughput sequencing technologies, the approximate string matching algorithms are used to identify similarities, molecular functions and abnormalities in DNA sequences. We consider a generalization of this problem, the fixed-length approximate string matching problem: given a text t, a pattern ρ and an integer ℓ, compute the optimal alignment of all substrings of ρ of length ℓ and a substring of t. We present a practical parallel algorithm of comparable simplicity that requires only time, where w is the word size of the machine (e.g. 32 or 64 in practice) and p the number of processors, by virtue of computing a bit representation of the relocatable dynamic programming matrix for the problem. Thus the algorithm's performance is independent of k and the alphabet size |Σ|. ",

author = "Iliopoulos, {Costas S.} and Laurent Mouchard and Pissis, {Solon P.}",

year = "2010",

doi = "10.3233/978-1-60750-530-3-150",

language = "Undefined/Unknown",

volume = "19",

pages = "150--157",

booktitle = "Proceedings of the International Conference on Parallel Computing (PARCO 2009)",

publisher = "IOS Press",

}

A Parallel Algorithm for the Fixed-length Approximate String Matching Problem for High Throughput Sequencing Technologies. / Iliopoulos, Costas S.; Mouchard, Laurent ; Pissis, Solon P.
Proceedings of the International Conference on Parallel Computing (PARCO 2009). Vol. 19 IOS Press, 2010. p. 150-157.

Research output: Chapter in Book/Report/Conference proceeding › Conference paper

TY - CHAP

T1 - A Parallel Algorithm for the Fixed-length Approximate String Matching Problem for High Throughput Sequencing Technologies

AU - Iliopoulos, Costas S.

AU - Mouchard, Laurent

AU - Pissis, Solon P.

PY - 2010

Y1 - 2010

N2 - The approximate string matching problem is to find all locations at which a query of length m matches a substring of a text of length n with k-or-fewer differences. Nowadays, with the advent of novel high throughput sequencing technologies, the approximate string matching algorithms are used to identify similarities, molecular functions and abnormalities in DNA sequences. We consider a generalization of this problem, the fixed-length approximate string matching problem: given a text t, a pattern ρ and an integer ℓ, compute the optimal alignment of all substrings of ρ of length ℓ and a substring of t. We present a practical parallel algorithm of comparable simplicity that requires only time, where w is the word size of the machine (e.g. 32 or 64 in practice) and p the number of processors, by virtue of computing a bit representation of the relocatable dynamic programming matrix for the problem. Thus the algorithm's performance is independent of k and the alphabet size |Σ|.

AB - The approximate string matching problem is to find all locations at which a query of length m matches a substring of a text of length n with k-or-fewer differences. Nowadays, with the advent of novel high throughput sequencing technologies, the approximate string matching algorithms are used to identify similarities, molecular functions and abnormalities in DNA sequences. We consider a generalization of this problem, the fixed-length approximate string matching problem: given a text t, a pattern ρ and an integer ℓ, compute the optimal alignment of all substrings of ρ of length ℓ and a substring of t. We present a practical parallel algorithm of comparable simplicity that requires only time, where w is the word size of the machine (e.g. 32 or 64 in practice) and p the number of processors, by virtue of computing a bit representation of the relocatable dynamic programming matrix for the problem. Thus the algorithm's performance is independent of k and the alphabet size |Σ|.

U2 - 10.3233/978-1-60750-530-3-150

DO - 10.3233/978-1-60750-530-3-150

M3 - Conference paper

VL - 19

SP - 150

EP - 157

BT - Proceedings of the International Conference on Parallel Computing (PARCO 2009)

PB - IOS Press

ER -