From attention to decoder-only transformers — the architecture behind GPT
Turning raw scores into probabilities
Turning words and categories into vectors
Query, key, value — the core of self-attention
Parallel attention heads learning different patterns
How transformers know word order
A visual guide to 'Attention Is All You Need'
The architecture behind GPT and modern LLMs