Lensa ML
Lensa ML

How Learning Rate Scheduling Works

Step 1 of 6

Constant LR Problems

Why a fixed learning rate fails

0501001500258EpochLossoptimal
Learning Rate0.50
LR
0.50
Converge
1 steps
Final Loss
0.69

High learning rate: loss oscillates as the model overshoots the minimum repeatedly.

Step 1 of 6: Constant LR Problems

Why a fixed learning rate fails

0501001500258EpochLossoptimal
Learning Rate0.50
LR
0.50
Converge
1 steps
Final Loss
0.69

High learning rate: loss oscillates as the model overshoots the minimum repeatedly.

Step 2 of 6: Step Decay

Reduce LR at fixed intervals

05010015000.050.1EpochLR
Decay Factor0.50
Step Size (epochs)30
Init LR
0.10
Decay
×0.50
Current LR
0.0063

Moderate step decay gives a nice staircase pattern. Each plateau lets the model settle before reducing the step size.

Step 3 of 6: Exponential Decay

Smooth continuous decrease

05010015000.050.1EpochLRe50e100
Decay Rate0.970
Decay Rate
0.970
LR @ 50
0.0218
LR @ 100
0.00476

Moderate decay rate produces a smooth curve that gradually reduces exploration over training.

Step 4 of 6: Cosine Annealing

Wave-like schedule with warm restarts

05010015020000.050.1EpochLRmin LR
T_max (cycle length)100
T_max
100
Min LR
0.001
LR @ 100
0.0010

Moderate cycle length creates a smooth annealing curve. The cosine shape spends more time at low LRs for fine-tuning.

Step 5 of 6: Warmup

Start slow, ramp up, then decay

05010015000.050.1EpochLRwarmup
Warmup Epochs10
Warmup
10 epochs
Peak LR
0.10
LR @ Peak
0.1000

Moderate warmup lets the model build reliable gradient statistics before applying the full learning rate. Standard for transformers.

Step 6 of 6: Cyclical LR

Periodic exploration and convergence

05010015020000.0450.09EpochLR
Cycle Length40
Amplitude0.080
Cycle
40 ep
Min LR
0.010
Max LR
0.090
Now
0.0900

Moderate cycles balance exploration and convergence. The triangular wave provides natural warmup and cooldown within each cycle.