GGTyper : genotyping complex structural variants using short-read sequencing data

dc.contributor.authorMirus, Tim
dc.contributor.authorLohmayer, Robert
dc.contributor.authorDöhring, Clementine
dc.contributor.authorHalldórsson, Bjarni V.
dc.contributor.authorKehr, Birte
dc.contributor.departmentDepartment of Engineering
dc.date.accessioned2026-10-07T14:01:01Z
dc.date.available2026-10-07T14:01:01Z
dc.date.issued2024-09-01
dc.descriptionPublisher Copyright: © The Author(s) 2024. Published by Oxford University Press.en
dc.description.abstractMotivation: Complex structural variants (SVs) are genomic rearrangements that involve multiple segments of DNA. They contribute to human diversity and have been shown to cause Mendelian disease. Nevertheless, our abilities to analyse complex SVs are very limited. As opposed to deletions and other canonical types of SVs, there are no established tools that have explicitly been designed for analysing complex SVs. Results: Here, we describe a new computational approach that we specifically designed for genotyping complex SVs in short-read sequenced genomes. Given a variant description, our approach computes genotype-specific probability distributions for observing aligned read pairs with a wide range of properties. Subsequently, these distributions can be used to efficiently determine the most likely genotype for any set of aligned read pairs observed in a sequenced genome. In addition, we use these distributions to compute a genotyping difficulty for a given variant, which predicts the amount of data needed to achieve a reliable call. Careful evaluation confirms that our approach outperforms other genotypers by making reliable genotype predictions across both simulated and real data. On up to 7829 human genomes, we achieve high concordance with population-genetic assumptions and expected inheritance patterns. On simulated data, we show that precision correlates well with our prediction of genotyping difficulty. This together with low memory and time requirements makes our approach well-suited for application in biomedical studies involving small to very large numbers of short-read sequenced genomes. Availability and implementation: Source code is available at https://github.com/kehrlab/Complex-SV-Genotyping.en
dc.description.versionPeer revieweden
dc.format.extent994086
dc.format.extentii11-ii19
dc.identifier.citationMirus, T, Lohmayer, R, Döhring, C, Halldórsson, B V & Kehr, B 2024, 'GGTyper : genotyping complex structural variants using short-read sequencing data', Bioinformatics, vol. 40, pp. ii11-ii19. https://doi.org/10.1093/bioinformatics/btae391en
dc.identifier.doi10.1093/bioinformatics/btae391
dc.identifier.issn1367-4803
dc.identifier.other251154351
dc.identifier.othere52b7220-822e-4695-b50e-1480c38b2540
dc.identifier.other85203210034
dc.identifier.other39230689
dc.identifier.urihttps://hdl.handle.net/20.500.11815/8571
dc.language.isoen
dc.relation.ispartofseriesBioinformatics; 40()en
dc.relation.urlhttps://www.scopus.com/pages/publications/85203210034en
dc.rightsinfo:eu-repo/semantics/openAccessen
dc.subjectStatistics and Probabilityen
dc.subjectBiochemistryen
dc.subjectMolecular Biologyen
dc.subjectComputer Science Applicationsen
dc.subjectComputational Theory and Mathematicsen
dc.subjectComputational Mathematicsen
dc.titleGGTyper : genotyping complex structural variants using short-read sequencing dataen
dc.type/dk/atira/pure/researchoutput/researchoutputtypes/contributiontojournal/articleen

Skrár

Original bundle

Niðurstöður 1 - 1 af 1
Nafn:
btae391.pdf
Stærð:
970.79 KB
Snið:
Adobe Portable Document Format