Initialization, gradient problems, overfitting, and regularization
Xavier, He, and why random weights matter
Why deep networks struggle and how to fix it
The bias-variance tradeoff and generalization
Dropout, L1, and L2 — keeping models honest