Class 10 · Unit 5: Natural language processing · Lesson 5.4 · 40 min
Bag of words and TF-IDF
Four slips from the suggestion box. Can a computer find the one word each slip is really about — without reading it?
Today you will: Build a bag of words by hand · Calculate TF-IDF · Let Python find each slip's key word
Story
The head teacher has a pile of suggestion slips and five minutes before the staff meeting.
Teacher: “I don't need every word. I need to know what each slip is about.”
Bilal counts the words on the first slip. Tea appears twice. On the third, tap appears twice.
Bilal: “Easy. The word that appears most.”
Hiba: “Then what about break? It's on two slips. And library, and books. Are those what every slip is about?”
Bilal looks again. "A word that's everywhere can't tell the slips apart."
Hiba: “So you want a word that's common on one slip, and rare on the others.”
Watch
Turning words into numbers — — scan code 10.5.4 in the printed book.
Warm up your fingers
Skill: a formula with LOG10 (5 min)
-
On typing.com: the symbols lessons.
-
Type three times, eyes on the screen:
=B2*LOG10(4/C2) -
Goal this term: 38–44 WPM at 95%.
On the laptop
You need: LibreOffice Calc · Thonny · programs/bow.py · datasets/suggestions.txt
Mission 1: Bag of words by hand (10 min, in pairs)
- ☐1
Here are the four slips after normalisation (Lesson 5.3):
- Slip 1: canteen, runs, tea, break, make, tea
- Slip 2: library, open, break, want, books
- Slip 3: tap, ground, broken, fix, tap
- Slip 4: keep, library, open, school, books
- ☐2
In Calc, list the vocabulary across row 1 — each word once. How many words is it?
- ☐3
Fill one row per slip with the count of each word. That is the document vector table.
Mission 2: TF-IDF (8 min)
- ☐4
Under the table, add a row for DF: in how many slips each word appears. Use
=COUNTIF(B2:B5;">0"). - ☐5
Predict: which word on slip 1 will have the highest TF-IDF? And on slip 3?
- ☐6
Make a second table: TF-IDF for every cell,
=B2*LOG10(4/B$7)(if DF is in row 7). Were you right?
Mission 3: Let Python find the key words (5 min)
- ☐7
Copy
bow.pynext tosuggestions.txt. Run it. It prints the vocabulary, the four document vectors, and each slip's key word. Does it agree with your table?
Finished early?
Slips 2 and 4 come out with key words want and keep. Add them — and make
— to the stop words in bow.py, and run it again. Are the new key words better?
No laptop today?
Build the document-vector table together on the board, one slip per row of desks. Then work out TF-IDF for tea, break and tap only, using log(4) = 0.602 and log(2) = 0.301.
Now you know
- Bag of words: a vocabulary of every unique word, and a document vector counting each word in each document. Word order is ignored.
- TF: how often a word appears in one document. DF: how many documents it appears in.
- TF-IDF = TF × log(N ÷ DF) , log base 10. High when a word is common in one document and rare in the rest.
- A word in every document scores 0, like a stop word — it can't tell them apart.
- Uses: classifying documents, finding topics, search, filtering stop words.
Debate it
Search engines rank pages partly with ideas like TF-IDF.
- If a page repeats a word a hundred times to rank higher, does TF-IDF catch it?
- Slips written in Kashmiri or Urdu would share no words with the English slips. What happens to them in this bag of words?
Check yourself — practice, not a test
1. In the bag of words model, the list of every unique word in the corpus is the ______.
2. A word appears in all 4 of 4 documents. Its TF-IDF is:
- ○ 4
- ○ 1
- ○ 0
- ○ 0.602
3. Tap appears twice in slip 3 and in no other slip. With 4 slips, its TF-IDF is:
- ○ 0.301
- ○ 0.602
- ○ 1.204
- ○ 2
4. In a bag of words, the order of the words matters.
- ○ True
- ○ False
5. Match each term to its meaning:
Term frequency · Document frequency
Match with: How many documents a word appears in · How often a word appears in one document
Remember
Common here, rare everywhere else — that's the word that matters.