Lensa ML
Lensa ML

Object Detection

Step 1 of 6

Classification vs Detection

From labels to bounding boxes

Number of Objects1
Cat (98.2%)Image Classification — one label per image
OBJECTS
1
MODE
Classify
OUTPUTS
1 label

Classification assigns a single label to the entire image. It answers "what is this?" but not "where is it?"

Step 1 of 6: Classification vs Detection

From labels to bounding boxes

Number of Objects1
Cat (98.2%)Image Classification — one label per image
OBJECTS
1
MODE
Classify
OUTPUTS
1 label

Classification assigns a single label to the entire image. It answers "what is this?" but not "where is it?"

Step 2 of 6: Anchor Boxes

Predefined reference shapes on a grid

Anchors per Cell1
12 total anchors across 4x3 grid
ANCHORS / CELL
1
GRID CELLS
12
TOTAL ANCHORS
12

With 1 square anchor per cell, we can only detect roughly square objects. Tall or wide objects may be missed.

Step 3 of 6: Bounding Box Regression

Refining anchor positions to match objects

Regression Offset0.00
Ground TruthAnchorPredicted
OFFSET
0.00
IoU
0.000
dx
0px

The predicted box is still close to the anchor. The network must learn larger offsets to reach the ground truth.

Step 4 of 6: Confidence & Class Scores

Filtering detections by confidence

Confidence Threshold0.30
Cat 0.95Dog 0.87Bird 0.62Car 0.41Tree 0.28Bike 0.15Person 0.73Threshold 0.30 — 5 kept, 2 removed
KEPT
5
REMOVED
2
PRECISION
0.72
RECALL
0.71

At threshold 0.30, weaker detections are filtered out. This balances precision (0.72) and recall (0.71).

Step 5 of 6: Non-Maximum Suppression

Removing duplicate detections

IoU Threshold0.50
0.950.820.680.910.74IoU(box1, box2) = 0.680 — threshold = 0.50
IoU THRESHOLD
0.50
INPUT BOXES
5
SURVIVING
2

NMS keeps the highest-confidence box and removes neighbors with IoU above 0.50. This removes duplicate detections while preserving distinct objects.

Step 6 of 6: Single-Stage vs Two-Stage

YOLO vs Faster R-CNN tradeoffs

Architecture Blend (YOLO ↔ Faster R-CNN)0.00
YOLO (Single-Stage)Single forward passFaster R-CNN (Two-Stage)RPNClassifierCatDogCarPropose → ClassifyFast / Less AccurateSlow / More Accurate
ARCHITECTURE
1-Stage
SPEED (FPS)
50
mAP
50.0%

Single-stage detectors like YOLO predict bounding boxes and classes in one pass over the grid. This yields high speed (~50 FPS) but may sacrifice accuracy on small objects.