Evaluation of whole-genome sequence data analysis approaches for short- and long-read sequencing of Mycobacterium tuberculosis

Peker, Nilay; Schuele, Leonard; Kok, Nienke; Terrazos, Miguel; Neuenschwander, Stefan M.; de Beer, Jessica; Akkerman, Onno; Peter, Silke; Ramette, Alban; Merker, Matthias; Niemann, Stefan; Couto, Natacha; Sinha, Bhanu; Rossen, John Wa

Published in

Microbiology Society, Microbial Genomics, 11(7), 2021

DOI: 10.1099/mgen.0.000695

Tools

Export citation

Search in Google Scholar

Evaluation of whole-genome sequence data analysis approaches for short- and long-read sequencing of Mycobacterium tuberculosis

Journal article published in 2021 by Nilay Peker

, Leonard Schuele, Nienke Kok, Miguel Terrazos

, Stefan M. Neuenschwander

, Jessica de Beer, Onno Akkerman

, Silke Peter, Alban Ramette

, Matthias Merker, Stefan Niemann, Natacha Couto

, Bhanu Sinha

, John Wa Rossen

This paper is made freely available by the publisher.

Full text: Download

Preprint: archiving allowed

Upload

Postprint: archiving allowed

Upload

Published version: archiving allowed

Upload

Policy details

Data provided by

Abstract

Whole-genome sequencing (WGS) of Mycobacterium tuberculosis (MTB) isolates can be used to get an accurate diagnosis, to guide clinical decision making, to control tuberculosis (TB) and for outbreak investigations. We evaluated the performance of long-read (LR) and/or short-read (SR) sequencing for anti-TB drug-resistance prediction using the TBProfiler and Mykrobe tools, the fraction of genome recovery, assembly accuracies and the robustness of two typing approaches based on core-genome SNP (cgSNP) typing and core-genome multi-locus sequence typing (cgMLST). Most of the discrepancies between phenotypic drug-susceptibility testing (DST) and drug-resistance prediction were observed for the first-line drugs rifampicin, isoniazid, pyrazinamide and ethambutol, mainly with LR sequence data. Resistance prediction to second-line drugs made by both TBProfiler and Mykrobe tools with SR- and LR-sequence data were in complete agreement with phenotypic DST except for one isolate. The SR assemblies were more accurate than the LR assemblies, having significantly (P<0.05) fewer indels and mismatches per 100 kbp. However, the hybrid and LR assemblies had slightly higher genome fractions. For LR assemblies, Canu followed by Racon, and Medaka polishing was the most accurate approach. The cgSNP approach, based on either reads or assemblies, was more robust than the cgMLST approach, especially for LR sequence data. In conclusion, anti-TB drug-resistance prediction, particularly with only LR sequence data, remains challenging, especially for first-line drugs. In addition, SR assemblies appear more accurate than LR ones, and reproducible phylogeny can be achieved using cgSNP approaches.

Published in

Links

Tools

Evaluation of whole-genome sequence data analysis approaches for short- and long-read sequencing of Mycobacterium tuberculosis

Abstract