Lensa ML
Lensa ML

Ensemble Methods & Stacking

Step 1 of 6

Why Ensembles?

Wisdom of diverse models

Number of models3
Model Predictions on 10 Test PointsS1S2S3S4S5S6S7S8S9S10Tree70%SVM90%k-NN80%ENS100%Accuracy Comparison70%Tree90%SVM80%k-NN100%Ens
MODELS
3
AVG INDIV
80%
ENSEMBLE
100%
IMPROVE
20.0%

Different models make different errors — combining 3 models via majority vote cancels out individual mistakes, boosting accuracy from 80% to 100%.

Step 1 of 6: Why Ensembles?

Wisdom of diverse models

Number of models3
Model Predictions on 10 Test PointsS1S2S3S4S5S6S7S8S9S10Tree70%SVM90%k-NN80%ENS100%Accuracy Comparison70%Tree90%SVM80%k-NN100%Ens
MODELS
3
AVG INDIV
80%
ENSEMBLE
100%
IMPROVE
20.0%

Different models make different errors — combining 3 models via majority vote cancels out individual mistakes, boosting accuracy from 80% to 100%.

Step 2 of 6: Hard Voting

Majority rules

Active models3
Hard Voting — Majority RulesS1S2S3S4S5S6S7S8Tree10000101SVM10100101k-NN11100101VOTE10100101TRUTH10100101Per-Model AccuracyTree88%SVM100%k-NN88%Ens100%
ACTIVE
3
S4 VOTES
0:3
WINNER
Class 0
ENS ACC
100%

Hard voting picks the class with the most votes. With 3 models, a prediction needs 2+ votes to win — simple but ignores confidence.

Step 3 of 6: Soft Voting

Weighted probabilities

Weight strategy (equal → accuracy-based)0.50
Sample to inspect2
Soft Voting — Sample S2 BreakdownEach model outputs a probability → weighted average decides the classTreew=0.89acc=78%0.34SVMw=0.91acc=82%0.08k-NNw=0.88acc=76%0.19AVG0.50.202Weighted Average Formula(0.89×0.34 + 0.91×0.08 + 0.88×0.19) / 2.68= 0.202 → Class 0 ✓All Samples0.87S10.20S20.85S30.21S40.17S50.67S6
STRATEGY
Blended
AVG PROB
0.202
PRED CLASS
0
ENS ACC
100%

Blended weights (50% accuracy-based). Each model's weight is a mix of equal (1.0) and its accuracy. This balances fairness with rewarding better models.

Step 4 of 6: Stacking: Level 0

Base model predictions

Fold1
Stacking Level 0 — Out-of-Fold PredictionsData split into 5 foldsVALFold 1TRAINFold 2TRAINFold 3TRAINFold 4TRAINFold 5Train on 4 folds → Predict on Fold 1Base Model Accuracy on Fold 1Tree81.5%SVM83.6%k-NN72.6%NB74.1%LR78.7%avg 78.1%
FOLD
1/5
TRAIN
40
VAL
10
AVG ACC
78.1%

Fold 1 is held out for validation. The 5 base models train on folds 2–5, then predict on fold 1. These out-of-fold predictions become input features for the meta-learner.

Step 5 of 6: Stacking: Level 1

The meta-learner

Regularization0.50
Stacking Level 1 — Meta-LearnerBase model outputs feed into a meta-learnerTree0.17SVM0.28k-NN0.15NB0.14LR0.26MetaLearner90.8%Stacked Acc0.17Tree0.28SVM0.15k-NN0.14NB0.26LRLearned Meta-Weights
#1 SVM
0.28
#2 LR
0.26
STACK ACC
90.8%
OVERFIT RISK
50%

The meta-learner balances model contributions — SVM and LR get the most weight. Regularization prevents any single model from dominating.

Step 6 of 6: Comparing Strategies

Voting vs. stacking

Comparing Ensemble Strategies60%70%80%90%100%82.2%Best Indiv85.6%Hard Vote87.3%Soft Vote91.2%StackingNoise Level15%Winner: Stacking at 91.2%
BEST INDIV
82.2%
HARD VOTE
85.6%
SOFT VOTE
87.3%
STACKING
91.2%

At low noise, all strategies perform similarly — simple voting may suffice.