Evaluating the necessity of PCR duplicate removal from next-generation sequencing data and a comparison of approaches

Research output: Contribution to journalArticlepeer-review

115 Scopus citations

Abstract

Background: Analyzing next-generation sequencing data is difficult because datasets are large, second generation sequencing platforms have high error rates, and because each position in the target genome (exome, transcriptome, etc.) is sequenced multiple times. Given these challenges, numerous bioinformatic algorithms have been developed to analyze these data. These algorithms aim to find an appropriate balance between data loss, errors, analysis time, and memory footprint. Typical analysis pipelines require multiple steps. If one or more of these steps is unnecessary, it would significantly decrease compute time and data manipulation to remove the step. One step in many pipelines is PCR duplicate removal, where PCR duplicates arise from multiple PCR products from the same template molecule binding on the flowcell. These are often removed because there is concern they can lead to false positive variant calls. Picard (MarkDuplicates) and SAMTools (rmdup) are the two main softwares used for PCR duplicate removal. Results: Approximately 92 % of the 17+ million variants called were called whether we removed duplicates with Picard or SAMTools, or left the PCR duplicates in the dataset. There were no significant differences between the unique variant sets when comparing the transition/transversion ratios (p = 1.0), percentage of novel variants (p = 0.99), average population frequencies (p = 0.99), and the percentage of protein-changing variants (p = 1.0). Results were similar for variants in the American College of Medical Genetics genes. Genotype concordance between NGS and SNP chips was above 99 % for all genotype groups (e.g., homozygous reference). Conclusions: Our results suggest that PCR duplicate removal has minimal effect on the accuracy of subsequent variant calls.

Original languageEnglish
Article number239
JournalBMC Bioinformatics
Volume17
DOIs
StatePublished - Jul 25 2016

Bibliographical note

Publisher Copyright:
© 2016 The Author(s).

Keywords

  • Next-Generation Sequencing
  • PCR duplicate removal
  • Picard
  • SAMTools

ASJC Scopus subject areas

  • Structural Biology
  • Biochemistry
  • Molecular Biology
  • Computer Science Applications
  • Applied Mathematics

Fingerprint

Dive into the research topics of 'Evaluating the necessity of PCR duplicate removal from next-generation sequencing data and a comparison of approaches'. Together they form a unique fingerprint.

Cite this