Optimal spliced alignments of short sequence reads

De Bona, F; Ossowski, S; Schneeberger, K; Rätsch, G

doi:10.1093/bioinformatics/btn300

Local TagsRelease HistoryDetailsSummary

Optimal spliced alignments of short sequence reads

De Bona, F., Ossowski, S., Schneeberger, K., & Rätsch, G. (2008). Optimal spliced alignments of short sequence reads. Bioinformatics, 24(16), i174-i180. doi:10.1093/bioinformatics/btn300.

Item is Released

show all hide all

Basic

show hide

Item Permalink: https://hdl.handle.net/21.11116/0000-000A-B031-7 Version Permalink: https://hdl.handle.net/21.11116/0000-000A-B032-6

Genre: Journal Article

Files

show Files

Locators

show

Creators

show

hide

Creators:
De Bona, F¹, Author
Ossowski, S, Author
Schneeberger, K, Author
Rätsch, G¹, Author

Affiliations:
1Rätsch Group, Friedrich Miescher Laboratory, Max Planck Society, ou_3378052

Content

show

hide

Free keywords: -

Abstract:

Motivation: Next generation sequencing technologies open exciting new possibilities for genome and transcriptome sequencing. While reads produced by these technologies are relatively short and error prone compared to the Sanger method their throughput is several magnitudes higher. To utilize such reads for transcriptome sequencing and gene structure identification, one needs to be able to accurately align the sequence reads over intron boundaries. This represents a significant challenge given their short length and inherent high error rate.

Results: We present a novel approach, called QPALMA, for computing accurate spliced alignments which takes advantage of the read's quality information as well as computational splice site predictions. Our method uses a training set of spliced reads with quality information and known alignments. It uses a large margin approach similar to support vector machines to estimate its parameters to maximize alignment accuracy. In computational experiments, we illustrate that the quality information as well as the splice site predictions help to improve the alignment quality. Finally, to facilitate mapping of massive amounts of sequencing data typically generated by the new technologies, we have combined our method with a fast mapping pipeline based on enhanced suffix arrays. Our algorithms were optimized and tested using reads produced with the Illumina Genome Analyzer for the model plant Arabidopsis thaliana.

Availability: Datasets for training and evaluation, additional results and a stand-alone alignment tool implemented in C++ and python are available at http://www.fml.mpg.de/raetsch/projects/qpalma.

Details

show

hide

Language(s):

Dates: Date issued: 2008-08

Publication Status: Issued

Pages: -

Publishing info: -

Table of Contents: -

Rev. Type: -

Identifiers: DOI: 10.1093/bioinformatics/btn300
PMID: 18689821

Degree: -

Event

show

Legal Case

show

Project information

show

Source 1

show

hide

Title: Bioinformatics

Source Genre: Journal

Creator(s):

Affiliations:

Publ. Info: Oxford : Oxford University Press

Pages: - Volume / Issue: 24 (16) Sequence Number: - Start / End Page: i174 - i180 Identifier: ISSN: 1367-4803
CoNE: https://pure.mpg.de/cone/journals/resource/954926969991