Loss functions, gradient descent, and the knobs that drive training
How networks measure their mistakes
Optimizing by following the slope downhill
Why and how to normalize features for faster convergence
The knobs you tune before training begins
How data is fed to the model in epochs, batches, and shuffles