Lensa ML
Lensa ML

Cross-Validation

Step 1 of 6

The Problem with One Split

Unstable estimates

Data Points — Train (gray) vs Test (green)Test Accuracy by Seed74%Seed
SEED
1
TRAIN ACC
87.8%
TEST ACC
73.9%
STD ACROSS
±5.9%

A single train/test split gives an unreliable estimate — the result depends on which data lands where.

Step 1 of 6: The Problem with One Split

Unstable estimates

Data Points — Train (gray) vs Test (green)Test Accuracy by Seed74%Seed
SEED
1
TRAIN ACC
87.8%
TEST ACC
73.9%
STD ACROSS
±5.9%

A single train/test split gives an unreliable estimate — the result depends on which data lands where.

Step 2 of 6: K-Fold Cross-Validation

Every point gets validated

5-Fold Split — 5 colored foldsF1F2F3F4F5Each Row = One Iteration (highlighted = validation)F1F2F3F4F5
K
5
FOLD SIZE
6
TRAIN/FOLD
24
VAL/FOLD
6

K-fold ensures every data point is used for both training and validation exactly once.

Step 3 of 6: Walking Through Folds

Rotate the validation set

Fold 1 / 5
Fold 1 as ValidationPer-Fold AccuracyFold 184.7%Fold 2Fold 3Fold 4Fold 5μ=84.7
FOLD
1/5
FOLD ACC
84.7%
MEAN
84.7%
STD
±0.0%

Each fold takes a turn as the validation set — we average all fold scores for a stable estimate.

Step 4 of 6: Stratified K-Fold

Preserving class balance

Stratified 5-Fold — Class Distribution per FoldClass 0 (20)Class 1 (10)F14/2F24/2F34/2F44/2F54/2ideal 33%
K
5
CLASS RATIO
33%
MAX DEV
0.0%
BALANCE
100%

Stratified folds maintain the same class ratio in every fold — critical for imbalanced datasets.

Step 5 of 6: Leave-One-Out

The extreme case

Test point: 1 / 15
Leave-One-Out: Point 1 held out123456789101112131415N iterations, each training on N−1 points1 / 15 iterationsCompute cost scales as O(N²) — 15² = 225 relative units
N
15
FOLDS
15
COST (N²)
225
VARIANCE
±0.63

LOO uses N folds — low bias but high variance and expensive for large datasets.

Step 6 of 6: Choosing k

Bias-variance of the estimator

CV Score vs k707580859095Score %251015202530k (number of folds)k=5k=1083.5±2.6
K
5
MEAN SCORE
83.5%
STD
±2.6%
COST
Low

Standard compromise — k=5 or k=10 balances bias, variance, and compute.