prior iterations, the neural nets used more direct non-linearities, including the tanh, however, the ReLU normally learns at a speedier rate in systems that have only a few layers, and this takes into consideration the preparation of a deep supervised system, without unsupervised pre-preparing.