Lensa ML
Lensa ML

Pooling & Downsampling

Step 1 of 6

Why Pool?

Reducing spatial dimensions

Feature Map Size8
Input: 8×82×2 poolPooled: 4×4
Original Cells
64
Pooled Cells
16
Reduction
75%

A 8×8 map has 64 cells. After 2×2 pooling it shrinks to 16 — a 75% reduction in downstream computation.

Step 1 of 6: Why Pool?

Reducing spatial dimensions

Feature Map Size8
Input: 8×82×2 poolPooled: 4×4
Original Cells
64
Pooled Cells
16
Reduction
75%

A 8×8 map has 64 cells. After 2×2 pooling it shrinks to 16 — a 75% reduction in downstream computation.

Step 2 of 6: Max Pooling

Keep the strongest signal

Pool Size
Window Position0
Input 8×81187139459324639225094477065301471551596854556513095627302614520Max Pool Output9869769785693967
Pool Size
2×2
Window Max
9
Output Size
4×4

The 2×2 window picks the maximum value (9) from each region. This preserves the strongest activation and discards weaker ones.

Step 3 of 6: Average Pooling

Smooth aggregation

Pool Size
Window Position0
Input1187139459324639225094477065301471551596854556513095627302614520Max Pool9869769785693967Avg Pool453.56.32.84445.34.84.35.31.35.34.33
Max Value
9
Avg Value
4
Difference
5.0

Max pooling picks the peak (9) while average pooling smooths to 4. Max is better for detecting sharp features; average preserves overall intensity.

Step 4 of 6: Stride & Overlap

Controlling output size

Stride2
Window Position0
Input 8×8 — stride 21187139459324639225094477065301471551596854556513095627302614520Output 4×49869769785693967
Stride
2
Output Size
4×4
Overlap?
No

Stride 2 with pool 2×2 means non-overlapping windows — the most common setting. Each cell is read exactly once, and spatial size halves.

Step 5 of 6: Global Average Pooling

Collapse to a vector

Number of Channels4
4 feature maps (4×4) → 4 scalars via GAP0.60.40.30.10.50.00.30.71.00.30.90.30.20.70.50.70.470.70.90.80.71.00.30.20.90.50.40.10.90.50.90.90.10.610.10.60.70.00.60.40.30.40.80.20.60.10.60.10.50.10.370.70.90.30.70.40.20.10.00.10.80.30.20.80.70.30.90.47FC → Classes
Channels
4
Input Params
64
After GAP
4
Reduction
16×

Each of the 4 feature maps is averaged into a single number, producing a 4-d vector. This eliminates spatial dimensions entirely — no fully-connected explosion.

Step 6 of 6: Pooling vs Strided Conv

Fixed vs learned downsampling

Compare
Input 6×6266666643096686021643411871394593246Max Pool2×2, s=2Output6698629390 paramsKernel 3×30.40.80.60.50.9-0.3-0.50.70.1Conv3×3, s=2Output1871719111191014144 params
Pool Params
0
Conv Params
144
Learnable?
No
Output Size
3×3

Max pooling is a fixed operation — it takes the maximum in each 2×2 window. Zero parameters, fast, and acts as regularization. But it cannot learn what information to keep.