Evaluation of genomic high-throughput sequencing data generated on Illumina 
HiSeq and Genome Analyzer systems

Minoche, A. E.; Dohm, J. C.; Himmelbauer, H.

Lokale TagsFreigabegeschichteDetailsÜbersicht

Evaluation of genomic high-throughput sequencing data generated on Illumina HiSeq and Genome Analyzer systems

Minoche, A. E., Dohm, J. C., & Himmelbauer, H. (2011). Evaluation of genomic high-throughput sequencing data generated on Illumina HiSeq and Genome Analyzer systems. Genome Biology, 12(11), R112. Retrieved from http://www.ncbi.nlm.nih.gov/pubmed/22067484 http://genomebiology.com/content/pdf/gb-2011-12-11-r112.pdf.

Item is Freigegeben

einblenden: alle ausblenden: alle

Basisdaten

einblenden: ausblenden:

Datensatz-Permalink: https://hdl.handle.net/11858/00-001M-0000-0010-789E-0 Versions-Permalink: https://hdl.handle.net/11858/00-001M-0000-0010-789F-E

Genre: Zeitschriftenartikel

ausblenden:

Urheber:
Minoche, A. E., Autor
Dohm, J. C.¹, Autor
Himmelbauer, H.¹, Autor

Affiliations:
1Dept. of Vertebrate Genomics (Head: Hans Lehrach), Max Planck Institute for Molecular Genetics, Max Planck Society, ou_1433550

Inhalt

einblenden:

ausblenden:

Schlagwörter: -

Zusammenfassung: ABSTRACT: BACKGROUND: The generation and analysis of high-throughput sequencing data are becoming a major component of many studies in molecular biology and medical research. Illumina's Genome Analyzer (GA) and HiSeq instruments are currently the most widely used sequencing devices. Here, we comprehensively evaluate properties of genomic HiSeq and GAIIx data derived from two plant genomes and one virus, with read lengths of 95 to 150 bases. RESULTS: We provide quantifications and evidence for GC bias, error rates, error sequence context, effects of quality filtering, and the reliability of quality values. By combining different filtering criteria we reduced error rates 7-fold at the expense of discarding 12.5% of alignable bases. While overall error rates are low in HiSeq data we observed regions of accumulated wrong base calls. Only 3% of all error positions accounted for 24.7% of all substitution errors. Analyzing the forward and reverse strands separately revealed error rates of up to 18.7%. Insertions and deletions occurred at very low rates on average but increased to up to 2% in homopolymers. A positive correlation between read coverage and GC content was found depending on the GC content range. CONCLUSIONS: The errors and biases we report have implications for the use and the interpretation of Illumina sequencing data. GAIIx and HiSeq data sets show slightly different error profiles. Quality filtering is essential to minimize downstream analysis artifacts. Supporting previous recommendations, the strand-specificity provides a criterion to distinguish sequencing errors from low abundance polymorphisms.

Details

einblenden:

ausblenden:

Sprache(n):

Datum: Erschienen: 2011

Publikationsstatus: Erschienen

Seiten: -

Ort, Verlag, Ausgabe: -

Inhaltsverzeichnis: -

Art der Begutachtung: -

Identifikatoren: eDoc: 584735
URI: http://www.ncbi.nlm.nih.gov/pubmed/22067484 http://genomebiology.com/content/pdf/gb-2011-12-11-r112.pdf

Art des Abschluß: -

Veranstaltung

einblenden:

Entscheidung

einblenden:

Projektinformation

einblenden:

Quelle 1

einblenden:

ausblenden:

Titel: Genome Biology

Genre der Quelle: Zeitschrift

Urheber:

Affiliations:

Ort, Verlag, Ausgabe: -

Seiten: - Band / Heft: 12 (11) Artikelnummer: - Start- / Endseite: R112 Identifikator: ISSN: 1465-6914 (Electronic) 1465-6906 (Linking)

Datensatz

Basisdaten

Dateien

Externe Referenzen

Urheber

Inhalt

Details

Veranstaltung

Entscheidung

Projektinformation

Quelle 1