Curated sequences that take you from the basics to the frontier — each step builds on the one before it.
Build up from a single neuron to training your first network
Master the training loop — optimizers, regularization, and normalization
From attention to decoder-only transformers — the architecture behind GPT
The calculus and linear algebra you need for deep learning
Distributions, Bayes, and information theory for ML
Shorter routes through one area — pick a subject and work straight through it.
From a single neuron to forward propagation and classification
7 stepsDerivatives, the chain rule, and the optimization math behind training
5 stepsVectors, tensors, eigenvalues, and the decompositions ML relies on
8 stepsDistributions, Bayes, estimation, and information theory end to end
8 stepsTrees, SVMs, clustering, and ensemble methods beyond neural nets
5 stepsLoss functions, gradient descent, and the knobs that drive training
6 stepsBackprop, activations, optimizers, and learning-rate scheduling
4 stepsInitialization, gradient problems, overfitting, and regularization
8 stepsValidation, cross-validation, the bias–variance tradeoff, and pitfalls
7 stepsTokenization, classic text features, and word embeddings
6 stepsClassification, seq2seq, NER, and language modeling with BERT
7 stepsAttention, multi-head and cross-attention, and the transformer stack
6 stepsCNNs, recurrent nets, autoencoders, GANs, and skip connections
4 stepsPooling, detection, segmentation, and data augmentation
4 stepsDiffusion models, RLHF, LoRA, and retrieval-augmented generation
5 stepsFeature pipelines, embeddings, transfer learning, and dimensionality reduction