Lensa ML
Lensa ML

Bagging & Random Forests

Step 1 of 6

Bootstrap Sampling

Drawing with replacement

Bootstrap sample seed1
Original Dataset1234567891011121314151617181920sample with replacementBootstrap Sample14317315111781667422092015181212Unique sampleDuplicate (drawn 2+ times)
SAMPLE SIZE
20
UNIQUE
15
DUPLICATES
5
% USED
75%

This sample uses most of the original data. On average, bootstrap sampling includes ~63.2% of unique points — the rest are duplicates drawn with replacement.

Step 1 of 6: Bootstrap Sampling

Drawing with replacement

Bootstrap sample seed1
Original Dataset1234567891011121314151617181920sample with replacementBootstrap Sample14317315111781667422092015181212Unique sampleDuplicate (drawn 2+ times)
SAMPLE SIZE
20
UNIQUE
15
DUPLICATES
5
% USED
75%

This sample uses most of the original data. On average, bootstrap sampling includes ~63.2% of unique points — the rest are duplicates drawn with replacement.

Step 2 of 6: Bagging

Many trees, one vote

Number of trees5
Individual Tree Predictions0T11T21T31T41T5Ensemble VoteClass 1votes: 1 for class 0, 4 for class 1
TREES
5
AVG ACCURACY
68.4%
ENSEMBLE ACC
77.7%
IMPROVEMENT
+9.3%

Bagging trains 5 trees on different bootstrap samples and combines their votes. The ensemble is more stable than any single tree.

Step 3 of 6: Random Feature Subsets

Decorrelating the trees

Max features per split3
Feature Selection per TreeHeightWeightAgeIncomeScoreTree 1Tree 2Tree 3Tree 4Tree Correlation0.56Diversity0.44
MAX FEATURES
3
CORRELATION
0.56
ENSEMBLE ACC
0.9%
DIVERSITY
0.44

Using 3 of 5 features balances individual tree quality with diversity. This is the sweet spot for Random Forests.

Step 4 of 6: Variance Reduction

Why ensembles win

Number of trees5
Predictions on a number lineTRUEensemble = 52.6Indiv. var: 75.5Ensemble var: 6.8Ensemble Variance vs # Trees# treesVariance0257911020304050
TREES
5
INDIV VARIANCE
75.5
ENSEMBLE VAR
6.8
REDUCTION ~1/N
1/5

Averaging 5 trees reduces variance substantially. Each additional tree helps smooth the prediction.

Step 5 of 6: Out-of-Bag Error

Free validation

Number of trees10
Training Matrixrows = data points, cols = treesT1T2T3T4T5T6T7T8T9T10x1x2x3x4x5x6x7x8x9x10In-bag (training)OOB (free validation)
TREES
10
OOB ERROR
30.0%
OOB RATE
33.0%
EXPECTED OOB
~37%

With 10 trees, most points have several OOB predictions, giving a reliable error estimate without a separate test set.

Step 6 of 6: Feature Importance

Which features matter?

Number of trees10
Feature ImportanceIncome39.4%Age23.5%Education18.7%Location12.8%Hours5.6%Dashed = true importance
TOP FEATURE
Income
IMPORTANCE
39.4%
TREES
10
STABILITY
0.90

With 10 trees, importance estimates are noisy. The ranking may shift between runs.