Class 10 · Unit 5: Natural language processing · Lesson 5.3 · 40 min
Text normalisation
“The TAP is broken!!!” and “the tap is broken” — to a computer, these share almost nothing. Until you clean them.
Today you will: Take text through all the steps of normalisation · Clean real suggestion slips in Python · Tell stemming from lemmatisation
Story
The suggestion box outside the staff room has been opened for the first time since spring. Hiba tips forty slips onto the table.
Hiba: “The TAP is broken!!! tap broken near ground. Please fix the tap. Tap. Broken. Again.”
Bilal: “Everyone's saying the same thing. Let's get the computer to count.”
He types them in and counts the words. TAP, tap, Tap. and tap. come out as four different words, and the most common word of all is the.
Hiba: “The computer thinks the biggest problem in this school is the.”
Watch
Cleaning text for a machine — — scan code 10.5.3 in the printed book.
Warm up your fingers
Skill: a list of short words (5 min)
-
On typing.com: the symbols lessons.
-
Type three times, eyes on the screen:
stop_words = ["the", "a", "is", "of", "to", "and"] -
Goal this term: 38–44 WPM at 95%.
On the laptop
You need: Thonny · programs/normalise.py · datasets/suggestions.txt
Mission 1: By hand (8 min, in pairs)
- ☐1
Take this slip: "The tap near the ground is broken. Please fix the tap."
- ☐2
Do each step on paper, one line per step: sentences → tokens → without stop words and full stops → lowercase. How many tokens are left at the end?
Mission 2: Let Python do it (10 min)
- ☐3
Copy
normalise.pynext tosuggestions.txtand open it in Thonny. Find the line that makes the text lowercase, the line that splits it into tokens, and the line that drops stop words. - ☐4
Predict: what will it print for the third slip? Compare with your paper answer.
- ☐5
Run it. Were you right?
- ☐6
Change it: add
"break"to the stop words. Predict, run, and decide: was that a good idea?
Mission 3: Stem or lemma? (5 min)
- ☐7
For each word, write its stem and its lemma: studies, running, better, flies, happily, children. Where do they differ?
Finished early?
Add a line that removes a final s from every token, as a crude stemmer. Run
it. Which words does it help — and which does it wreck?
No laptop today?
Everyone gets one slip on paper and normalises it by hand, step by step, then reads out their final tokens. The class writes the shared tokens on the board and tallies them — the counting that Lesson 5.4 does in Python.
Now you know
- The corpus is all the text you're working with. Normalisation cleans it so a machine can count and compare it.
- Segment, then tokenise: corpus → sentences → tokens (every word, number and symbol).
- Remove stop words very common words that carry little meaning — and any symbols or numbers the job doesn't need.
- One case: lowercase everything, so "Tap" and "tap" are the same word.
- Stemming chops endings fast (studies → studi). Lemmatisation gives a real word (studies → study).
Debate it
A stop-word list is a choice somebody makes. For some jobs, removing "not" would reverse the meaning of a sentence.
- Find a sentence where removing a "stop word" changes its meaning completely.
- Who should choose the stop words for a system that reads complaints from the public?
Check yourself — practice, not a test
1. Put CBSE's steps of text normalisation in order:
- 1. Tokenisation
- 2. Converting to a common case
- 3. Removing stop words, special characters and numbers
- 4. Sentence segmentation
- 5. Stemming or lemmatisation
2. Which of these is most likely a stop word?
- ○ tap
- ○ broken
- ○ the
- ○ library
3. The whole collection of text being processed is called the ______.
4. Match each word to what it becomes:
studies by stemming · studies by lemmatisation
Match with: study · studi5. Lemmatisation always produces a real word, but it takes longer than stemming.
- ○ True
- ○ False
Remember
Clean the text before you count it — or you'll count "the".