Skip to main navigation Skip to search Skip to main content

Systematic analysis of dark and camouflaged genes reveals disease-relevant genes hiding in plain sight

  • Mark T.W. Ebbert
  • , Tanner D. Jensen
  • , Karen Jansen-West
  • , Jonathon P. Sens
  • , Joseph S. Reddy
  • , Perry G. Ridge
  • , John S.K. Kauwe
  • , Veronique Belzil
  • , Luc Pregent
  • , Minerva M. Carrasquillo
  • , Dirk Keene
  • , Eric Larson
  • , Paul Crane
  • , Yan W. Asmann
  • , Nilufer Ertekin-Taner
  • , Steven G. Younkin
  • , Owen A. Ross
  • , Rosa Rademakers
  • , Leonard Petrucelli
  • , John D. Fryer

Research output: Contribution to journalArticlepeer-review

136 Scopus citations

Abstract

Background: The human genome contains "dark" gene regions that cannot be adequately assembled or aligned using standard short-read sequencing technologies, preventing researchers from identifying mutations within these gene regions that may be relevant to human disease. Here, we identify regions with few mappable reads that we call dark by depth, and others that have ambiguous alignment, called camouflaged. We assess how well long-read or linked-read technologies resolve these regions. Results: Based on standard whole-genome Illumina sequencing data, we identify 36,794 dark regions in 6054 gene bodies from pathways important to human health, development, and reproduction. Of these gene bodies, 8.7% are completely dark and 35.2% are ≥ 5% dark. We identify dark regions that are present in protein-coding exons across 748 genes. Linked-read or long-read sequencing technologies from 10x Genomics, PacBio, and Oxford Nanopore Technologies reduce dark protein-coding regions to approximately 50.5%, 35.6%, and 9.6%, respectively. We present an algorithm to resolve most camouflaged regions and apply it to the Alzheimer's Disease Sequencing Project. We rescue a rare ten-nucleotide frameshift deletion in CR1, a top Alzheimer's disease gene, found in disease cases but not in controls. Conclusions: While we could not formally assess the association of the CR1 frameshift mutation with Alzheimer's disease due to insufficient sample-size, we believe it merits investigating in a larger cohort. There remain thousands of potentially important genomic regions overlooked by short-read sequencing that are largely resolved by long-read technologies.

Original languageEnglish
Article number97
JournalGenome Biology
Volume20
Issue number1
DOIs
StatePublished - May 20 2019

Bibliographical note

Publisher Copyright:
© 2019 The Author(s).

Funding

This work was supported by the PhRMA Foundation [RSGTMT17 to M.E.]; the Ed and Ethel Moore Alzheimer's Disease Research Program of Florida Department of Health [8AZ10 and 9AZ08 to M.E., and 6AZ06 to J.F.]; the Muscular Dystrophy Association (M.E.); the National Institutes of Health [NS094137 to J.F., AG047327 to J. F, AG049992 to J.F., NS097261 to R.R., NS097273 to L.P., NS084528 to L.P., NS084974 to L.P., NS099114 to L.P., NS088689 to L.P., NS093865 to L.P.]; Department of Defense [ALSRP AL130125 to L.P.]; Mayo Clinic Foundation (L.P. and J.F.); Mayo Clinic Center for Individualized Medicine (L.P. and J.F.); Amyotrophic Lateral Sclerosis Association (M.E., L.P.); Robert Packard Center for ALS Research at Johns Hopkins (L.P.) Target ALS (L.P.); Association for Frontotemporal Degeneration (L.P.); GHR Foundation (J.F.); and the Mayo Clinic Gerstner Family Career Development Award (J.F.).

FundersFunder number
National Institutes of Health (NIH)
European Commission
Association for Frontotemporal Degeneration
Muscular Dystrophy Association
National Institute on AgingR01AG033193, R01AG023629, R21AG047327, U01AG057659, R01AG054076, U01AG049508, U24AG021886, U01AG046139, U01AG049505, U01AG049506, U01AG016976, U01AG049507, U54AG052427, U01AG052411, R01AG061796, U01AG052410, R01AG033040, U01AG032984, UF1AG047133, R01AG049607, RF1AG051504, U01AG052409, U24AG041689, R03AG049992, R01AG020098, R01AG015928
National Human Genome Research InstituteU54HG003273, U54HG003079, U54HG003067
Austrian Science Fund/FWFI 904, P 13180
Institute of Neurological Disorders and Stroke National Advisory Neurological Disorders and Stroke CouncilR01NS093865, R01NS017950, R01NS094137, R01NS088689, P01NS099114, R35NS097261, R21NS084528, R35NS097273, P01NS084974
???publication-publication-funding-organisation-not-added???047.017.043
National Heart, Lung, and Blood Institute (NHLBI)U01HL080295, R01HL085083, RC2HL102419, R01HL105756, U01HL096899, U01HL096812, U01HL096814, U01HL096902, U01HL096917, R01HL070825, U01HL130114
Seventh Framework Programme201413

    UN SDGs

    This output contributes to the following UN Sustainable Development Goals (SDGs)

    1. SDG 3 - Good Health and Well-being
      SDG 3 Good Health and Well-being

    Keywords

    • 10x Genomics
    • APOE
    • Alzheimer's Disease Sequencing Project (ADSP)
    • CR1
    • Camouflaged genes
    • Dark genes
    • Long-read sequencing
    • Oxford Nanopore Technologies (ONT)
    • Pacific Biosciences (PacBio)

    ASJC Scopus subject areas

    • Ecology, Evolution, Behavior and Systematics
    • Genetics
    • Cell Biology

    Fingerprint

    Dive into the research topics of 'Systematic analysis of dark and camouflaged genes reveals disease-relevant genes hiding in plain sight'. Together they form a unique fingerprint.

    Cite this