Lensa ML
Lensa ML

Feature Selection

Step 1 of 6

Why Select Features?

Less is often more

FEATURES
5
TRAIN ACC
83.7%
TEST ACC
79.2%
GAP
4.5%
50%60%70%80%90%100%15101520# FeaturesTrainTest
# Features: 5

Adding informative features improves the model. Each feature carries useful signal about the target.

Step 1 of 6: Why Select Features?

Less is often more

FEATURES
5
TRAIN ACC
83.7%
TEST ACC
79.2%
GAP
4.5%
50%60%70%80%90%100%15101520# FeaturesTrainTest
# Features: 5

Adding informative features improves the model. Each feature carries useful signal about the target.

Step 2 of 6: Filter Methods

Score features independently

THRESHOLD
0.30
KEPT
5
REMOVED
15
EST. ACC
92.0%
Age0.90Income0.80Score0.69Height0.55Weight0.42Rand20.17Noise10.16Rand50.16Rand80.16ID0.16Rand30.15Rand10.14Rand100.14Color0.13Rand60.12Noise30.11Noise20.09Rand90.09Rand70.08Rand40.05threshold=0.30
Correlation threshold: 0.30

Filter methods rank features by statistical measures like correlation — fast and scalable but they ignore feature interactions.

Step 3 of 6: Forward Selection

Add one at a time

ROUND
0
SELECTED
0
ACCURACY
50.0%
NEXT BEST
Age
FeaturesAge0.89Income0.74Score0.61Height0.45Weight0.33Color0.08ID0.06Noise10.05Noise20.04Noise30.04Accuracy by Round050.0%12345678910
0/10

Forward selection greedily adds the feature that most improves the model at each step. Press Play or drag the slider to start.

Step 4 of 6: L1 Regularization

Let the model decide

λ
0.004
NON-ZERO
20
ZEROED
0
ACCURACY
80.5%
Age1.206Income1.085Score0.980Height0.850Weight0.753Rand50.154Color0.143Rand80.142Rand70.129Noise10.128Rand30.125Rand40.085Noise20.079Noise30.075Rand90.067ID0.057Rand20.054Rand10.050Rand60.047Rand100.033
L1 penalty λ: 0.0040

Low penalty — all features retained. The model uses every available signal, including noise.

Step 5 of 6: Feature Importance

Tree-based scores

TOP FEATURE
Age
TOP SCORE
0.858
# TREES
20
STABILITY
±5.6%
Age0.858Income0.736Score0.591Height0.395Weight0.351ID0.098Color0.088Noise30.084Noise10.070Rand20.057Rand40.053Rand10.046Rand90.035Rand80.032Rand100.027Rand60.020Noise20.018Rand30.009Rand50.005Rand70.005
# Trees: 20

Tree-based models naturally rank features by how much they reduce impurity across all splits. Importance scores are becoming more stable as the ensemble grows.

Step 6 of 6: Impact on Generalization

Selected vs. all features

# FEATURES
5
TRAIN ACC
69.0%
TEST ACC
92.1%
OPTIMAL #
5
50%60%70%80%90%100%optimal15101520# Top Features KeptTrainTest
# Top features: 5

The optimal feature count balances information gain against the curse of dimensionality. You're at the sweet spot! Maximum test accuracy with minimal overfitting.