From Ravi et al. - “Deep Learning for Health Informatics”.
Deep Neural Network

Description
- General deep framework usually used for classification or regression.
- Made of many hidden layers (more than 2).
- Allows complex (non-linear) hypotheses to be expressed.
Pros
- Widely used with successes in many areas.
Cons
- Training is not trivial because once the errors are back-propagated to the first few layers, they become miniscule (vanishing gradient problem).
- The learning process can be very slow.
Deep Autoencoder

Description
- Proposed in Hinton et al. - “Reducing the Dimensionality of Data with Neural Networks” and is mainly designed for feature extraction or dimensionality reduction.
- Has the same number of input and output nodes.
- Aims to recreate the input vector.
- Unsupervised learning method.
Pros
- Does not require labelled data for training.
- Many variations have been proposed to make the representation more robust: Sparse AutEnc. (Poultney et al. - “Efficient learning of sparse representations with an energy-based model”), Denoising AutEnc. (Vincent et al. - “Extracting and composing robust features with denoising autoencoders”), Contractive AutEnc. (Rifai et al. - “Contractive auto-encoders: Explicit invariance during feature extraction”), Convolutional AutEnc. (Masci et al. - “Stacked convo- lutional auto-encoders for hierarchical feature extraction”).
Cons
- Requires a pre-training stage.
- Training can also suffer from vanishing of the errors.
Deep Belief Network

Description
- Proposed in Hinton et al. - “A fast learning algorithm for deep belief nets” - is a composition of RBM where each sub-network’s hidden layer serves as the visible layer for the next.
- Has undirected connections just at the top two layers.
- Allows unsupervised and supervised training of the network.
Pros
- Proposes a layer-by-layer greedy learning strategy to initialize the network.
- Inferences tractable maximizing the likelihood directly.
Cons
- Training procedure is computationally expensive due to the initialization process and sampling.
Deep Boltzmann Machine

Description
- Proposed in Salakhutdinov and Hinton - “Deep boltzmann machines”, it is another approach based on the Boltzmann family.
- Processes undirected connections (conditionally independent) between all layers of the network.
- Uses a stochastic maximum likelihood (Younes - “On the convergence of markovian stochastic algorithms with rapidly decreasing ergodicity rates”) algorithm to maximize the lower bound of the likelihood.
Pros
- Incoroprates top-down feedback for more robust inferences with ambiguous inputs.
Cons
- Time complexity for the inference is higher than DBN.
- Optimization of the parameters is not practical for large datasets.
Recurrent Neural Network

Description
- Proposed in Williams and Zipser - “A learning algorithm for continually running fully recurrent neural networks”, is a NN capable of analyzing a stream of data.
- Useful in applications where the output depends on the previous computations.
- Shares the same weights across all steps.
Pros
- Can memorize sequential events.
- Can model time dependencies.
- Has shown great success in many Natural Language Processing applications.
Cons
- Learning issues are frequent due to the vanishing gradient and exploding gradient problems.
Convolutional Neural Network

Description
- Proposed in LeCun et al. - “Gradient-based learning applied to document recognition”, it is well suited for 2D data such as images.
- Every hidden convolutional filter transforms its input to a 3D output volume of neuron activations.
- Inspired by the neurobiological model of the visual cortex (Hubel and Wiesel - “Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex”).
Pros
- Few neuron connections required with respect to a typical NN.
- Many variants have been proposed: AlexNet (Krizhevsky et al. - “Imagenet classification with deep convolutional neural networks”), Clarifai (Zeiler and Fergus - “Visualizing and understanding convolutional networks”), and GoogLeNet (Szegedy et al. - Going deeper with convolutions”).
Cons
- It may require many layers to find an entire hierarchy of visual features.
- It usually requires a large dataset of labelled images.