Identifying and removing haplotypic duplication in primary genome assemblies

Guan, Dengfeng; McCarthy, Shane A.; Wood, Jonathan; Howe, Kerstin; Wang, Yadong; Durbin, Richard

Published in

Oxford University Press, Bioinformatics, 9(36), p. 2896-2898, 2020

DOI: 10.1093/bioinformatics/btaa025

Tools

Export citation

Search in Google Scholar

Identifying and removing haplotypic duplication in primary genome assemblies

Journal article published in 2019 by Dengfeng Guan

, Shane A. McCarthy

, Jonathan Wood

, Kerstin Howe

, Yadong Wang, Richard Durbin

This paper is made freely available by the publisher.

Full text: Download

Preprint: archiving allowed

Upload

Postprint: archiving restricted

Upload

Published version: archiving forbidden

Policy details

Data provided by

Abstract

Abstract Motivation Rapid development in long-read sequencing and scaffolding technologies is accelerating the production of reference-quality assemblies for large eukaryotic genomes. However, haplotype divergence in regions of high heterozygosity often results in assemblers creating two copies rather than one copy of a region, leading to breaks in contiguity and compromising downstream steps such as gene annotation. Several tools have been developed to resolve this problem. However, they either focus only on removing contained duplicate regions, also known as haplotigs, or fail to use all the relevant information and hence make errors. Results Here we present a novel tool, purge_dups, that uses sequence similarity and read depth to automatically identify and remove both haplotigs and heterozygous overlaps. In comparison with current tools, we demonstrate that purge_dups can reduce heterozygous duplication and increase assembly continuity while maintaining completeness of the primary assembly. Moreover, purge_dups is fully automatic and can easily be integrated into assembly pipelines. Availability and implementation The source code is written in C and is available at https://github.com/dfguan/purge_dups. Supplementary information Supplementary data are available at Bioinformatics online.

Published in

Links

Tools

Identifying and removing haplotypic duplication in primary genome assemblies

Abstract