Class 10 · Unit 5: Natural language processing · Lesson 5.5 · 40 min
Sentiment analysis
“Not bad for the price.” Happy or unhappy? Your brain knew instantly. The computer got it backwards.
Today you will: Score reviews with a word list in Python · Find exactly where the scorer fails · Measure it with Unit 3's tools
Story
Hiba's uncle rents a houseboat on Dal Lake to tourists in summer. A booking website sends him a report every month: "Guest sentiment: 72% positive."
Uncle: “What does that mean?”
Bilal: “A computer read your reviews, and decided which were happy.”
Her uncle reads one aloud. "Not bad for the price."
Uncle: “That's a good review. That man came back twice.”
Bilal checks the report. The computer has put it under negative.
Hiba: “It saw the word bad. It didn't see the word not.”
Watch
Reading feelings from text — — scan code 10.5.5 in the printed book.
Warm up your fingers
Skill: two word lists (5 min)
-
On typing.com: the symbols lessons.
-
Type three times, eyes on the screen:
positive_words = ["clean", "friendly", "delicious", "peaceful"] -
Goal this term: 38–44 WPM at 95%.
On the laptop
You need: Thonny · programs/sentiment.py · datasets/reviews.csv · LibreOffice Calc
Mission 1: Score by hand (5 min, in pairs)
- ☐1
Open
reviews.csvin Calc: 16 houseboat reviews, each with its true label. - ☐2
Predict: read the reviews without the labels. Which three will a word-counting computer get wrong? Write their numbers.
Mission 2: Run the scorer (10 min)
- ☐3
Copy
sentiment.pynext toreviews.csv. Read the two word lists and the scoring loop. - ☐4
Run it. It prints every review it gets wrong, and its score out of 16. How many of your three predictions were among them?
- ☐5
Sort its mistakes into negation, sarcasm, unknown word and other.
Mission 3: Improve it — honestly (8 min)
- ☐6
Make one improvement: add zabardast and wah to the positive words, or handle not by flipping the score of the next word.
- ☐7
Run it again. What's the new score out of 16?
- ☐8
Build a confusion matrix for the new version, with negative as the positive class. What is its recall for unhappy guests?
Finished early?
Write three new reviews your improved scorer will still get wrong. Swap them with another pair and see if their scorer gets them right.
No laptop today?
Your teacher reads the 16 reviews. The class scores each one aloud with the two word lists, adding +1 and −1 on the board. Then compare with the true labels and sort the mistakes into the four kinds.
Now you know
- Sentiment analysis finds the feeling in text — positive, negative or neutral.
- A word-list scorer adds +1 for positive words and −1 for negative ones. Simple, fast, and easy to read.
- It fails on negation, sarcasm, mixed feelings and words it doesn't know including our own languages.
- It is a classifier, so measure it like one: accuracy, a confusion matrix, recall for the class that matters.
- Tools: Python with NLTK or spaCy (code); Orange (no code). Bigger models do better and still miss local meaning.
Debate it
The booking website shows other tourists the uncle's "72% positive" score.
- If the scorer misreads "Not bad for the price" and "Zabardast!", who pays for its mistakes?
- Should a business be judged by a number a machine produced from its reviews? What should the website tell people about how the number was made?
Check yourself — practice, not a test
1. "The room was not clean." A word-list scorer counts clean as +1. This failure is called:
- ○ Sarcasm
- ○ Negation
- ○ An unknown word
- ○ Tokenisation
2. A scorer gets 11 of 16 reviews right. Its accuracy is about:
- ○ 11%
- ○ 16%
- ○ 69%
- ○ 91%
3. A sentiment scorer is a kind of classifier, and can be evaluated with a confusion matrix.
- ○ True
- ○ False
4. Which would trick a simple word-list scorer? (choose all that apply)
- ○ "Nothing was wonderful about it."
- ○ "Great. Another cold night."
- ○ "Wah! Zabardast!"
- ○ "The room was dirty."
5. Finding whether a piece of text is positive or negative is called ______ analysis.
Remember
Words have feelings in context. A word list only sees the words.