Bagging & Random Forests
Bootstrap Sampling
Drawing with replacement
This sample uses most of the original data. On average, bootstrap sampling includes ~63.2% of unique points — the rest are duplicates drawn with replacement.
Step 1 of 6: Bootstrap Sampling
Drawing with replacement
This sample uses most of the original data. On average, bootstrap sampling includes ~63.2% of unique points — the rest are duplicates drawn with replacement.
Step 2 of 6: Bagging
Many trees, one vote
Bagging trains 5 trees on different bootstrap samples and combines their votes. The ensemble is more stable than any single tree.
Step 3 of 6: Random Feature Subsets
Decorrelating the trees
Using 3 of 5 features balances individual tree quality with diversity. This is the sweet spot for Random Forests.
Step 4 of 6: Variance Reduction
Why ensembles win
Averaging 5 trees reduces variance substantially. Each additional tree helps smooth the prediction.
Step 5 of 6: Out-of-Bag Error
Free validation
With 10 trees, most points have several OOB predictions, giving a reliable error estimate without a separate test set.
Step 6 of 6: Feature Importance
Which features matter?
With 10 trees, importance estimates are noisy. The ranking may shift between runs.