Towards an unbiased characterization of genetic polymorphism

Igolkina, AA; Vorbrugg, S; Rabanal, FA; Liu, H-J; Ashkenazy, H; Kornienko, AE; Fitz, J; Collenberg, M; Kubica, C; Mollá Morales, A; Jaegle, B; Wrightsman, T; Voloshin, V; Llaca, V; Nizhynska, V; Reichardt, I; Lanz, C; Bemm, F; Flood, BJ; Nemomissa, S; Hancock, A; Guo, Y-L; Kersey, P; Weigel, D; Nordborg, M

doi:10.1101/2024.05.30.596703

Towards an unbiased characterization of genetic polymorphism

Igolkina, A., Vorbrugg, S., Rabanal, F., Liu, H.-J., Ashkenazy, H., Kornienko, A., Fitz, J., Collenberg, M., Kubica, C., Mollá Morales, A., Jaegle, B., Wrightsman, T., Voloshin, V., Llaca, V., Nizhynska, V., Reichardt, I., Lanz, C., Bemm, F., Flood, B., Nemomissa, S., Hancock, A., Guo, Y.-L., Kersey, P., Weigel, D., & Nordborg, M. (submitted). Towards an unbiased characterization of genetic polymorphism.

Item is 公開

表示: 全項目非表示: 全項目

基本情報

表示: 非表示:

アイテムのパーマリンク: https://hdl.handle.net/21.11116/0000-000F-6088-B 版のパーマリンク: https://hdl.handle.net/21.11116/0000-000F-6089-A

資料種別: Preprint

ファイル

表示: ファイル

作成者

表示:

非表示:

作成者:
Igolkina, AA, 著者
Vorbrugg, S¹, 著者
Rabanal, FA¹, 著者
Liu, H-J, 著者
Ashkenazy, H¹, 著者
Kornienko, AE, 著者
Fitz, J¹, 著者
Collenberg, M¹, 著者
Kubica, C¹, 著者
Mollá Morales, A, 著者
Jaegle, B, 著者
Wrightsman, T¹, 著者
Voloshin, V, 著者
Llaca, V, 著者
Nizhynska, V, 著者
Reichardt, I, 著者
Lanz, C¹, 著者
Bemm, F¹, 著者
Flood, BJ, 著者
Nemomissa, S, 著者
Hancock, A, 著者Guo, Y-L, 著者Kersey, P, 著者Weigel, D¹, 著者                Nordborg, M, 著者全て表示

所属:
1Department Molecular Biology, Max Planck Institute for Biology Tübingen, Max Planck Society, ou_3371687

内容説明

表示:

非表示:

キーワード: -

要旨: Our view of genetic polymorphism is shaped by methods that provide a limited and reference-biased picture. Long-read sequencing technologies, which are starting to provide nearly complete genome sequences for population samples, should solve the problem—except that characterizing and making sense of non-SNP variation is difficult even with perfect sequence data. Here, we analyze 27 genomes of Arabidopsis thaliana in an attempt to address these issues, and illustrate what can be learned by analyzing whole-genome polymorphism data in an unbiased manner. Estimated genome sizes range from 135 to 155 Mb, with differences almost entirely due to centromeric and rDNA repeats. The completely assembled chromosome arms comprise roughly 120 Mb in all accessions, but are full of structural variants, many of which are caused by insertions of transposable elements (TEs) and subsequent partial deletions of such insertions. Even with only 27 accessions, a pan-genome coordinate system that includes the resulting variation ends up being 40% larger than the size of any one genome. Our analysis reveals an incompletely annotated mobile-ome: our ability to predict what is actually moving is poor, and we detect several novel TE families. In contrast to this, the genic portion, or “gene-ome”, is highly conserved. By annotating each genome using accession-specific transcriptome data, we find that 13% of all genes are segregating in our 27 accessions, but that most of these are transcriptionally silenced. Finally, we show that with short-read data we previously massively underestimated genetic variation of all kinds, including SNPs—mostly in regions where short reads could not be mapped reliably, but also where reads were mapped incorrectly. We demonstrate that SNP-calling errors can be biased by the choice of reference genome, and that RNA-seq and BS-seq results can be strongly affected by mapping reads to a reference genome rather than to the genome of the assayed individual. In conclusion, while whole-genome polymorphism data pose tremendous analytical challenges, they will ultimately revolutionize our understanding of genome evolution.

資料詳細

表示:

非表示:

言語:

日付: 投稿: 2024-05

出版の状態: 投稿済み

ページ: -

出版情報: -

目次: -

査読: -

識別子（DOI, ISBNなど）: DOI: 10.1101/2024.05.30.596703

学位: -

アイテム詳細

基本情報

ファイル

関連URL

作成者

内容説明

資料詳細

関連イベント

訴訟

Project information

出版物