Lensa ML
Lensa ML

Bag of Words

Step 1 of 6

Building a Vocabulary

From corpus to word list

Minimum frequency threshold1
Corpus5 docs · 30 tokensD0"the cat sat on the mat"6D1"the dog ran in the park"6D2"a cat and a dog played"6D3"the park had a big mat"6D4"a dog sat on a mat"6Word Frequenciesacross all docsa5the5dog3mat3cat2on2park2sat2and1big1had1in1played1ran1min freq = 1
VOCAB SIZE
14
REMOVED
0
TOTAL TOKENS
30

Minimum frequency = 1: every word in the corpus is included. The vocabulary has all unique words — common ones like 'the' (6×) and rare ones like 'big' (1×).

Step 1 of 6: Building a Vocabulary

From corpus to word list

Minimum frequency threshold1
Corpus5 docs · 30 tokensD0"the cat sat on the mat"6D1"the dog ran in the park"6D2"a cat and a dog played"6D3"the park had a big mat"6D4"a dog sat on a mat"6Word Frequenciesacross all docsa5the5dog3mat3cat2on2park2sat2and1big1had1in1played1ran1min freq = 1
VOCAB SIZE
14
REMOVED
0
TOTAL TOKENS
30

Minimum frequency = 1: every word in the corpus is included. The vocabulary has all unique words — common ones like 'the' (6×) and rare ones like 'big' (1×).

Step 2 of 6: One-Hot Encoding

One word, one dimension

Select word0
Vocabularyindex → one-hot vectoridxwordone-hot vector0a100000000000001the010000000000002dog001000000000003mat000100000000004cat000010000000005on000001000000006park000000100000007sat000000010000008and000000001000009big0000000001000010had0000000000100011in0000000000010012played0000000000001013ran00000000000001Vector dimension = vocabulary size = 14
WORD
a
INDEX
0
DIMENSIONS
14
NON-ZERO
1

"a" is at index 0 → its one-hot vector has a 1 at position 0 and 0s everywhere else. Common words get low indices. Every word occupies exactly one dimension — no notion of similarity between words yet.

Step 3 of 6: Count Vectors

Summing one-hots into a document vector

Select document0
Doc 0"the cat sat on the mat"theidx 1catidx 4satidx 7onidx 5theidx 1matidx 3Σ one-hots → count vectorCount Vectora0the2dog0mat1cat1on1park0sat1and0big0had0in0played0ran0
DOCUMENT
Doc 0
TOKENS
6
UNIQUE
5
SPARSITY
64%

Doc 0 has 6 tokens mapping to 5 unique words. Words like "the" appear 2× — the count vector captures frequency, not position.

Step 4 of 6: Document-Term Matrix

The full picture — all documents, all terms

Vocabulary size10
Document-Term Matrix5 × 10athedogmatcatonparksatandbigD00201110100D10210001000D22010100010D31101001001D42011010100Rows = documents (5) · Columns = terms (10) · Cells = word countsNon-zero: 22 / 50 (56% sparse)
MATRIX
5×10
NON-ZERO
22
SPARSITY
56%
TERMS
10

The document-term matrix is the full BoW picture: every document is a row, every term is a column. With 10 terms, the matrix is 56% sparse — most cells are 0. Real corpora are 99%+ sparse.

Step 5 of 6: Binary vs. Count vs. Normalized

Three flavors of BoW

CountRaw word frequencies — longer docs naturally have higher countsathedogmatcatonparksatandbigD00201110100D10210001000D22010100010D31101001001D42011010100Value range: 0–2 · Non-zero: 22 / 50
MODE
Count
VALUE RANGE
0-2
NON-ZERO
22

Count BoW: raw word counts. 'the' appears 2× in Doc 0, so its cell is 2. Preserves frequency, but longer documents naturally have higher counts — unfair for comparison.

Step 6 of 6: Limitations

What BoW loses

Order invariance: different order, same BoWSentence:dogbitesmanMeaning: "A dog bites a man"BoW vector:bites1dog1man1⚠ Same vector for ALL orderings!Sparsity problem (Doc 0, full vocab):non-zerozeros (64%)Key limitations:⚠ No word order — 'dog bites man' = 'man bites dog'⚠ High dimensionality — one dimension per vocabulary word⚠ Extreme sparsity — most entries are zero⚠ No semantic similarity — 'happy' and 'glad' are orthogonal
ORDERING
#1
BOW VECTOR
[1,1,1]
SPARSITY
64%

"Dog bites man" — a normal headline. The BoW vector is [1,1,1]. Now slide to reorder the words and watch: the meaning changes completely, but the BoW vector stays the same.