When does Vanishing Gradients occur?


Vanishing Gradient occurs when the derivative or slope becomes smaller as we go backward layer by layer during backpropagation.

When weights update is small the training time takes long; this may even bring to a complete halt the neural network training.

Vanishing Gradient: with sigmoid and tanh activation fn, as the derivatives of sigmoid and tanh activation functions are 0 – 0.25 and 0 – 1.

Hence the updated weight values are small, with new weight values much like old weight values. This leads to Vanishing Gradient problem.

Avoid this by using ReLU activation fn, b/c the gradient is 0 for negatives with zero input –> 1 for positive input.

Exploding Gradient is the opposite of the vanishing gradients. This is a ‘weights’ problem, not an ‘activation fn. With high weight values, derivatives are higher so weight considerable relative to the older weight.

Thus the gradient can never converge. You may get only oscillation around minima, yet not get to a global minimal point.

Summary: Backpropagation in the deep neural networks – the Vanishing Gradient problem results from the sigmoid and tan activation fn while the Exploding Gradient problem occurs due to large weights.

by: Roderick (Rodd) M.

Ads Blocker Image Powered by Code Help Pro

Ads Blocker Detected!!!

We have detected that you are using extensions to block ads. Please support us by disabling these ads blocker.

Powered By
Best Wordpress Adblock Detecting Plugin | CHP Adblock