Lensa ML
Lensa ML

Vanishing & Exploding Gradients

Step 1 of 6

Gradient Flow

Forward & backward pass through a real network

Hidden layers: 3(2–3–3–3–1)
inputh1h2h3output
avg |∂L/∂a|1e-4in9e-4h10.006h20.063h31.000out
Click a neuron to see its forward & backward calculations

Architecture

2–3–3–3–1

Avg Input Grad

1.1e-4

Output Act

0.413

With 3 hidden layers, notice the gradient magnitude weakening toward the input. Each sigmoid squeezes the gradient by up to 4×.

Step 1 of 6: Gradient Flow

Forward & backward pass through a real network

Hidden layers: 3(2–3–3–3–1)
inputh1h2h3output
avg |∂L/∂a|1e-4in9e-4h10.006h20.063h31.000out
Click a neuron to see its forward & backward calculations

Architecture

2–3–3–3–1

Avg Input Grad

1.1e-4

Output Act

0.413

With 3 hidden layers, notice the gradient magnitude weakening toward the input. Each sigmoid squeezes the gradient by up to 4×.

Step 2 of 6: Exploding Gradients

Linear network where gradients blow up backward

Hidden layers: 3|Weight scale: 1.5×(2–3–3–3–1)
inputh1h2h3output
avg |∂L/∂a|0.3in0.4h10.4h20.5h31.0out
Click a neuron to see its forward & backward calculations

Architecture

2–3–3–3–1

Avg Input Grad

0.35

Status

Mild

With weight scale 1.5× and 3 hidden layers, gradients are still manageable. But increase either and watch them blow up.

Step 3 of 6: Sigmoid vs ReLU

Why activation choice decides gradient fate

Input x: 0.00|Layers: 5
-4-2240.251.0σ'(x) — max 0.25ReLU'(x) — 0 or 1Gradient after 5 layers: (deriv)^5σ0.0010ReLU0.0

σ'(x)

0.250

ReLU'(x)

0

σ after 5L

0.0010

ReLU after 5L

0.0

At x=0, sigmoid has its best derivative (0.25) and ReLU is right at its kink. Even at sigmoid's peak, it still shrinks gradients 4× per layer.

Step 4 of 6: Training Impact

How gradient problems stall real training

Depth: 5 layers|LR: 1.0(2–3–3–3–1)
0.50epoch (real gradient descent, 2–3–3–3–1 sigmoid network, 4 training samples)2L5Lavg |Δw|5e-5h10.0003h20.0030h30.0701out

Final Loss (5L)

0.1327

Final Loss (2L)

0.0239

L1 avg |Δw|

5.2e-5

At 5 layers, training is slower but converges. The first hidden layer's weight update (5.2e-5) is much smaller than the last layer's (0.0701) — vanishing in action.

Step 5 of 6: Skip Connections

How identity shortcuts rescue gradient flow

Hidden layers: 5(2–3–3–3–3–3–1)
inputh1h2h3h4h5output
avg |grad|no skipskip4e-60.01in5e-50.08h14e-40.01h20.000.08h30.030.03h40.090.08h51.01.0out
Click a neuron to see its forward & backward calculations

Input Grad (no skip)

3.9e-6

Input Grad (skip)

0.0121

Improvement

3e+3×

Without skip connections, the gradient must pass through every sigmoid (σ' ≤ 0.25). After 5 layers, the input gradient is 3.9e-6. Toggle "Skip Connections" to see how residual blocks rescue gradient flow.

Step 6 of 6: Gradient Clipping

The safety net for exploding gradients

Max norm: 5.0
Before3.2-7.15.5-2.88.3-6.0After1.1-2.51.9-1.02.9-2.1

Original ‖g‖

14.27

Threshold

5.0

Clipped ‖g‖

5.00

Gradient norm 14.3 exceeds threshold 5.0. All components are scaled by 0.350 to bring the norm down, preserving direction but limiting magnitude.