TY - JOUR T1 - Nucleotide Archival Format (NAF) enables efficient lossless reference-free compression of DNA sequences JF - bioRxiv DO - 10.1101/501130 SP - 501130 AU - Kirill Kryukov AU - Mahoko Takahashi Ueda AU - So Nakagawa AU - Tadashi Imanishi Y1 - 2018/01/01 UR - http://biorxiv.org/content/early/2018/12/19/501130.abstract N2 - Summary: DNA sequence databases use compression such as gzip to reduce the required storage space and network transmission time. We describe Nucleotide Archival Format (NAF) - a new file format for lossless reference-free compression of FASTA and FASTQ-formatted nucleotide sequences. NAF compression ratio is comparable to the best DNA compressors, while providing dramatically faster decompression. We compared our format with DNA compressors: DELIMINATE and MFCompress, and with general purpose compressors: gzip, bzip2, xz, brotli, and zstd.Availability and implementation NAF compressor and decompressor, as well as format specification are available at https://github.com/KirillKryukov/naf. Format specification is in public domain. Compressor and decompressor are open source under the zlib/libpng license, free for nearly any use.Contact kkryukov{at}gmail.com ER -