Derivatives, the chain rule, and the optimization math behind training
The calculus behind gradient descent
How derivatives compose through layers
Extend single-variable derivatives to functions of multiple variables and gradient vectors
The math behind cross-entropy, softmax, and information theory
Partial derivatives, gradients, and the Jacobian — calculus for vectors
Why some optimization problems are harder than others
Lagrange multipliers — optimization with constraints