TY - JOUR T1 - Reverse-Complement Equivariant Networks for DNA Sequences JF - bioRxiv DO - 10.1101/2021.06.03.446953 SP - 2021.06.03.446953 AU - Vincent Mallet AU - Jean-Philippe Vert Y1 - 2021/01/01 UR - http://biorxiv.org/content/early/2021/10/23/2021.06.03.446953.abstract N2 - As DNA sequencing technologies keep improving in scale and cost, there is a growing need to develop machine learning models to analyze DNA sequences, e.g., to decipher regulatory signals from DNA fragments bound by a particular protein of interest. As a double helix made of two complementary strands, a DNA fragment can be sequenced as two equivalent, so-called Reverse Complement (RC) sequences of nucleotides. To take into account this inherent symmetry of the data in machine learning models can facilitate learning. In this sense, several authors have recently proposed particular RC-equivariant convolutional neural networks (CNNs). However, it remains unknown whether other RC-equivariant architectures exist, which could potentially increase the set of basic models adapted to DNA sequences for practitioners. Here, we close this gap by characterizing the set of all linear RC-equivariant layers, and show in particular that new architectures exist beyond the ones already explored. We further discuss RC-equivariant pointwise nonlinearities adapted to different architectures, as well as RC-equivariant embeddings of k-mers as an alternative to one-hot encoding of nucleotides. We show experimentally that the new architectures can outperform existing ones.Competing Interest StatementThe authors have declared no competing interest. ER -