to encourage generalization of the network, we should shuffle training data so batches do not contain the same examples between epochs
to encourage generalization of the network, we should shuffle training data so batches do not contain the same examples between epochs