The Juicebox Assembly Tools module facilitates de novo assembly of mammalian genomes with chromosome-length scaffolds for under $1000

Abstract
Hi-C contact maps are valuable for genome assembly (Lieberman-Aiden, van Berkum et al. 2009; Burton et al. 2013; Dudchenko et al. 2017). Recently, we developed Juicebox, a system for the visual exploration of Hi-C data (Durand, Robinson et al. 2016), and 3D-DNA, an automated pipeline for using Hi-C data to assemble genomes (Dudchenko et al. 2017). Here, we introduce “Assembly Tools,” a new module for Juicebox, which provides a point-and-click interface for using Hi-C heatmaps to identify and correct errors in a genome assembly. Together, 3D-DNA and the Juicebox Assembly Tools greatly reduce the cost of accurately assembling complex eukaryotic genomes. To illustrate, we generated de novo assemblies with chromosome-length scaffolds for three mammals: the wombat, Vombatus ursinus (3.3Gb), the Virginia opossum, Didelphis virginiana (3.3Gb), and the raccoon, Procyon lotor (2.5Gb). The only inputs for each assembly were Illumina reads from a short insert DNA-Seq library (300 million Illumina reads, maximum length 2x150 bases) and an in situ Hi-C library (100 million Illumina reads, maximum read length 2x150 bases), which cost <$1000.
Subject Area
- Biochemistry (13912)
- Bioengineering (10594)
- Bioinformatics (33689)
- Biophysics (17353)
- Cancer Biology (14414)
- Cell Biology (20424)
- Clinical Trials (138)
- Developmental Biology (11004)
- Ecology (16236)
- Epidemiology (2067)
- Evolutionary Biology (20551)
- Genetics (13532)
- Genomics (18832)
- Immunology (13973)
- Microbiology (32612)
- Molecular Biology (13568)
- Neuroscience (71074)
- Paleontology (533)
- Pathology (2226)
- Pharmacology and Toxicology (3785)
- Physiology (5973)
- Plant Biology (12176)
- Synthetic Biology (3409)
- Systems Biology (8255)
- Zoology (1878)