Attention, multi-head and cross-attention, and the transformer stack
Turning raw scores into probabilities
Query, key, value — the core of self-attention
Parallel attention heads learning different patterns
How transformers know word order
Attending across sequences and modalities
A visual guide to 'Attention Is All You Need'
The architecture behind GPT and modern LLMs