How Entropy & Cross-Entropy Work
Surprise
Unlikely events carry more information
P(event)
0.50
Surprise
1.00 bits
Likely events are not very surprising — we already expected them.
Step 1 of 6: Surprise
Unlikely events carry more information
P(event)
0.50
Surprise
1.00 bits
Likely events are not very surprising — we already expected them.
Step 2 of 6: Entropy
The average surprise of a distribution
H(X) = −Σ P(X)·log₂(P(X))
P(X)
0.50
1 − P(X)
0.50
ENTROPY
H(X) = 1.000
Maximum entropy! When both outcomes are equally likely, uncertainty is highest — 1 full bit.
Step 3 of 6: Distributions & Entropy
Shape determines uncertainty
Entropy H
2.000 bits
Max H
2.000 bits
Near-uniform distribution — maximum uncertainty. All outcomes are equally likely.
Step 4 of 6: Cross-Entropy
Measuring the mismatch between distributions
H(p)
1.157 bits
H(p,q)
1.446 bits
KL gap
0.290 bits
Decent match — there is a small KL divergence gap. Cross-entropy equals entropy plus the KL divergence penalty.
Step 5 of 6: Cross-Entropy as Loss
The ML connection — why this is the loss function
ŷ
0.50
Loss
0.693 nats
∂L/∂ŷ
-2.0
The model is fairly wrong — loss is rising and the gradient pushes strongly to fix the prediction.
Step 6 of 6: Why Minimize Cross-Entropy
Watch training converge to the true distribution
Epoch
0
H(p,q)
1.589 bits
Accuracy
46.0%
Before training begins, the model predicts uniformly — cross-entropy starts high. Press Play or step through epochs.