Lensa ML
Lensa ML

Image Segmentation

Step 1 of 6

Segmentation vs Detection

From image-level labels to pixel-level masks

Granularity0.00
"Outdoor Scene"classification
Level
classification
Granularity
0%
Output
1 label

Classification assigns one label to the entire image. Quick, but no spatial information — we only know WHAT is in the scene, not WHERE.

Step 1 of 6: Segmentation vs Detection

From image-level labels to pixel-level masks

Granularity0.00
"Outdoor Scene"classification
Level
classification
Granularity
0%
Output
1 label

Classification assigns one label to the entire image. Quick, but no spatial information — we only know WHAT is in the scene, not WHERE.

Step 2 of 6: Semantic Segmentation

Every pixel gets a class label — sky, road, car, person

Number of classes3
SkyTreeRoad
Classes
3
Total Pixels
100
Largest Class
46 px

With 3 classes, the model must learn distinct boundaries between each region. Per-class counts: Sky=46, Tree=17, Road=37.

Step 3 of 6: Encoder-Decoder Architecture

Compress then expand — the U-shaped backbone of segmentation

Layer depth2
Encoder3ch64chDecoder64ch3chResolution: 60px → bottleneck → 60px
Depth
2
Bottleneck
64ch
Downsample

Depth 2: one downsampling step halves the spatial resolution while doubling channels. The decoder upsamples back to the original size.

Step 4 of 6: Skip Connections

U-Net's key insight: bridge encoder details to the decoder

Skip connections1
EncoderDecoderskipskipskipIoU: 0.87 | Boundary: 92%
Skip Connections
ON
IoU Score
0.87
Boundary Acc
92%

Skip connections (U-Net style) concatenate encoder features with decoder features at matching resolutions. This preserves fine spatial details — boundaries are sharp and IoU jumps from 0.64 to 0.87.

Step 5 of 6: Instance Segmentation

Same class, different objects — each gets its own mask

Number of instances2
Semantic: all same color#1#2Obj 1Obj 2
Instances
2
Total Pixels
18
Class
Same
Unique Masks
2

2 objects of the same class. Semantic segmentation colors both identically. Instance segmentation assigns each a unique ID and mask — critical for counting and tracking.

Step 6 of 6: Upsampling Methods

Nearest, bilinear, or learned — how to grow feature maps back

Upsampling method0
Input 3×30.20.80.30.90.40.70.10.60.5Output 6×60.200.200.800.800.300.300.200.200.800.800.300.300.900.900.400.400.700.700.900.900.400.400.700.700.100.100.600.600.500.500.100.100.600.600.500.50Nearest Neighbor
Method
Nearest
Unique Values
9
Smoothness
Low

Nearest neighbor: each output pixel copies the closest input pixel. Fast and parameter-free, but produces blocky artifacts — visible as 2×2 blocks of identical values.