Lensa ML
Lensa ML

Class Imbalance & Resampling

Step 1 of 6

What is Class Imbalance?

When one class dominates

Minority class %10%
ratio9.0:1majorityminority
MAJORITY
90
MINORITY
10
RATIO
9.0:1

Moderately imbalanced — the minority class is under-represented and most classifiers will struggle.

Step 1 of 6: What is Class Imbalance?

When one class dominates

Minority class %10%
ratio9.0:1majorityminority
MAJORITY
90
MINORITY
10
RATIO
9.0:1

Moderately imbalanced — the minority class is under-represented and most classifiers will struggle.

Step 2 of 6: The Accuracy Paradox

99% accurate, 0% useful

Minority class %5%
"Predict all majority" classifierPred: MajorityPred: MinorityActual MajActual MinTN190FP0FN10TP0Accuracy: 95.0%But recall on minority = 0% — it misses every case!
ACCURACY
95.0%
PRECISION
N/A
RECALL
0%
F1 SCORE
N/A

A model that always predicts the majority class gets 95.0% accuracy but zero recall on the minority — it never catches fraud, disease, or rare events. Accuracy alone is misleading when classes are imbalanced.

Step 3 of 6: Random Oversampling

Duplicate the minority

Oversampling ratio1x
majority (n=50)minority originalduplicated
ORIGINAL MIN
10
AFTER OVERSAMPLE
10
DATASET SIZE
60
NEW RATIO
5.0:1

No oversampling applied — the minority class remains under-represented.

Step 4 of 6: SMOTE

Synthetic minority samples

Synthetic points0
original minoritysynthetic (SMOTE)majority
ORIGINAL
8
SYNTHETIC
0
TOTAL MINORITY
8
AVG NN DIST

No synthetic points yet. Drag the slider to generate SMOTE samples between nearest-neighbor pairs.

Step 5 of 6: Class Weights

Penalize majority errors less

Minority class weight1x
boundarypredict majoritypredict minority
WEIGHT
1x
BOUNDARY
200
PRECISION
100%
RECALL
100%

Equal weights — the decision boundary favors the majority class, missing many minority examples.

Step 6 of 6: Better Metrics

Beyond accuracy

Classification threshold0.50
RecallPrecision0.250.250.500.500.750.751.001.00Precision–Recall Curve
THRESHOLD
0.50
PRECISION
70%
RECALL
51%
F1 SCORE
59%

Near-optimal threshold — best F1 of 59% balances precision and recall. This is the sweet spot.