Machine Learning Algorithms for Polymorphism Detection

Zeller, G; Schweikert, G; Clark, R; Ossowski, S; Shin, P; Frazer, K; Ecker, J; Weigel, D; Schölkopf, B; Rätsch, G

Local TagsRelease HistoryDetailsSummary

Machine Learning Algorithms for Polymorphism Detection

Zeller, G., Schweikert, G., Clark, R., Ossowski, S., Shin, P., Frazer, K., et al. (2007). Machine Learning Algorithms for Polymorphism Detection. Talk presented at NIPS 2007 Workshop on Machine Learning in Computational Biology (MLCB 2007). Whistler, Canada. 2007-12-07 - 2007-12-08.

Item is Released

show all hide all

Basic

show hide

Item Permalink: https://hdl.handle.net/21.11116/0000-000A-7D07-3 Version Permalink: https://hdl.handle.net/21.11116/0000-000A-7D08-2

Genre: Talk

Files

show Files

Locators

show

hide

Locator:
http://media.nips.cc/Conferences/2007/NIPS-2007-Workshop-Book.pdf (Abstract) Open Access status unknown

Description:
-

OA-Status:

Creators

show

hide

Creators:
Zeller, G, Author
Schweikert, G, Author
Clark, R, Author
Ossowski, S, Author
Shin, P, Author
Frazer, K, Author
Ecker, J, Author
Weigel, D, Author
Schölkopf, B, Author
Rätsch, G¹, Author

Affiliations:
1Rätsch Group, Friedrich Miescher Laboratory, Max Planck Society, Max-Planck-Ring 9, 72076 Tübingen, DE, ou_3378052

Content

show

hide

Free keywords: -

Abstract: As extensive studies of natural variation require the identification of sequence differences among complete genomes, there exists a high demand for precise high-throughput sequencing techniques. While high-density oligo-nucleotide arrays are capable of rapid and comparatively cheap genomic scans, the resulting data is typically much noisier than dideoxy sequencing data. Therefore algorithmic approaches for the accurate identification of sequence polymorphisms from oligo-nucleotide array data remain a challenge [Gresham et al., 2006]. We present machine learning based methods tackling the problem of identifying Single Nucleotide Polymorphisms (SNPs) as well as deletions and highly polymorphic regions. Here we describe polymorphism discovery in 20 wild strains of the model plant Arabidopsis thaliana, which has a genome of about 125 Mb. A huge set of array hybridization data comprising nearly 19.2 billion measurements has been collected at Perlegen Sciences Inc. (four 25 nt probes for each base on each genomic strand and strain.

Details

show

hide

Language(s):

Dates: Published Online: 2007-12

Publication Status: Published online

Pages: -

Publishing info: -

Table of Contents: -

Rev. Type: -

Identifiers: BibTex Citekey: 5402

Degree: -

Event

show

hide

Title: NIPS 2007 Workshop on Machine Learning in Computational Biology (MLCB 2007)

Place of Event: Whistler, Canada

Start-/End Date: 2007-12-07 - 2007-12-08

Invited: Yes

Legal Case

show

Project information

show

Source

show