RT Journal Article SR Electronic T1 species2vec: A novel method for species representation JF bioRxiv FD Cold Spring Harbor Laboratory SP 461996 DO 10.1101/461996 A1 Angelov, Boyan YR 2019 UL http://biorxiv.org/content/early/2019/12/22/461996.abstract AB Word embeddings are omnipresent in Natural Language Processing (NLP) tasks. The same technology which defines words by their context can also define biological species. This study showcases this new method - species embedding (species2vec). By proximity sorting of 6761594 mammal observations from the whole world (2862 different species), we are able to create a training corpus for the skip-gram model. The resulting species embeddings are tested in an environmental classification task. The classifier performance confirms the utility of those embeddings in preserving the relationships between species, and also being representative of species consortia in an environment.