Vanishing & Exploding Gradients
Gradient Flow
Forward & backward pass through a real network
Architecture
2–3–3–3–1
Avg Input Grad
1.1e-4
Output Act
0.413
With 3 hidden layers, notice the gradient magnitude weakening toward the input. Each sigmoid squeezes the gradient by up to 4×.
Step 1 of 6: Gradient Flow
Forward & backward pass through a real network
Architecture
2–3–3–3–1
Avg Input Grad
1.1e-4
Output Act
0.413
With 3 hidden layers, notice the gradient magnitude weakening toward the input. Each sigmoid squeezes the gradient by up to 4×.
Step 2 of 6: Exploding Gradients
Linear network where gradients blow up backward
Architecture
2–3–3–3–1
Avg Input Grad
0.35
Status
Mild
With weight scale 1.5× and 3 hidden layers, gradients are still manageable. But increase either and watch them blow up.
Step 3 of 6: Sigmoid vs ReLU
Why activation choice decides gradient fate
σ'(x)
0.250
ReLU'(x)
0
σ after 5L
0.0010
ReLU after 5L
0.0
At x=0, sigmoid has its best derivative (0.25) and ReLU is right at its kink. Even at sigmoid's peak, it still shrinks gradients 4× per layer.
Step 4 of 6: Training Impact
How gradient problems stall real training
Final Loss (5L)
0.1327
Final Loss (2L)
0.0239
L1 avg |Δw|
5.2e-5
At 5 layers, training is slower but converges. The first hidden layer's weight update (5.2e-5) is much smaller than the last layer's (0.0701) — vanishing in action.
Step 5 of 6: Skip Connections
How identity shortcuts rescue gradient flow
Input Grad (no skip)
3.9e-6
Input Grad (skip)
0.0121
Improvement
3e+3×
Without skip connections, the gradient must pass through every sigmoid (σ' ≤ 0.25). After 5 layers, the input gradient is 3.9e-6. Toggle "Skip Connections" to see how residual blocks rescue gradient flow.
Step 6 of 6: Gradient Clipping
The safety net for exploding gradients
Original ‖g‖
14.27
Threshold
5.0
Clipped ‖g‖
5.00
Gradient norm 14.3 exceeds threshold 5.0. All components are scaled by 0.350 to bring the norm down, preserving direction but limiting magnitude.