Class 10 · Unit 3: Evaluating models · Lesson 3.4 · 40 min
Precision, recall and F1
A flood warning and a spam filter make opposite mistakes on purpose. Which one should never miss?
Today you will: Calculate precision, recall and F1 · Count a confusion matrix in Python · Choose the metric a problem needs
Story
In August the Jhelum rises fast. The district is trying out a flood-warning model: every evening it says flood or no flood for each village by the river.
Bilal: “It cried wolf four times last month. People have stopped listening.”
Hiba: “And how many floods did it miss?”
Bilal: “None.”
Hiba: “Then I'd keep it. A false alarm costs a night's sleep. A missed flood costs a house.”
Bilal thinks about his own phone. Last week its spam filter hid a message from school.
Bilal: “So my spam filter should be the other way round. It had better be sure before it hides anything.”
Watch
Two questions, two metrics — — scan code 10.3.4 in the printed book.
Warm up your fingers
Skill: typing Python with quotes and brackets (5 min)
-
On typing.com: the symbols lessons.
-
Type three times, eyes on the screen:
if actual == "scab" and predicted == "scab": -
Two equals signs to compare, one to store. Goal this term: 38–44 WPM at 95%.
On the laptop
You need: Thonny · programs/metrics.py · datasets/leaf-predictions.csv
Mission 1: By hand first (7 min, in pairs)
- ☐1
Copy the leaf model's matrix from Lesson 3.3: TP 11, FN 4, FP 6, TN 39.
- ☐2
Work out precision, recall and F1 on paper or in Calc. Round to two decimals.
Mission 2: Let Python count (12 min)
- ☐3
Copy
metrics.pyinto thedatasetsfolder, next toleaf-predictions.csv, and open it in Thonny. Read it. A row is a dictionary:row["actual"]is that leaf's actual class. - ☐4
Predict: what will the first line it prints say? Write it down.
- ☐5
Run it. Does it match your prediction — and your answers from Part 1?
- ☐6
Change it: add a line that also prints the accuracy. Predict the number, then run it.
Mission 3: Choose the metric (6 min)
- ☐7
For each problem, write precision, recall or F1, and the reason in five words:
- A flood warning for river villages
- A spam filter on a school email account
- A test that screens children for a vision problem
- A system that decides which days are safe to launch a satellite
- A bank flagging payments that might be fraud
- A filter keeping unsafe videos away from young children
Finished early?
Change the program so the positive class is healthy instead of scab. What happens to precision and recall — and does the change make sense?
No laptop today?
Do Part 1 on the board. Then your teacher reads metrics.py aloud, line by line, while the class
keeps four tallies — TP, FN, FP, TN — for the first ten rows of the CSV. Do Part 3 as a vote.
Now you know
- Precision = TP ÷ (TP + FP). How much of what it flagged was real. Matters when a false alarm is costly.
- Recall = TP ÷ (TP + FN). How much of what was real it caught. Matters when a miss is costly.
- F1 = 2 × P × R ÷ (P + R). One balanced number when both matter.
- Ask one question to choose: which does more harm here — a false alarm, or a miss?
- A program can count it for you, reading the CSV row by row. It will count exactly what you tell it to — so choose the positive class carefully.
Debate it
The flood model is tuned for high recall, so it cries wolf. People have started ignoring it.
- If people ignore the warnings, is high recall still helping them? What would you change?
- Who should decide how many false alarms are acceptable — the model's makers, or the villages?
Check yourself — practice, not a test
1. A disease test must not miss sick people. Which metric matters most?
- ○ Precision
- ○ Recall
- ○ Accuracy
- ○ Error rate
2. Precision = TP ÷ (TP + ______).
3. TP = 11, FP = 6. What is the precision?
- ○ 0.35
- ○ 0.65
- ○ 0.73
- ○ 0.83
4. F1 score combines precision and recall into one number.
- ○ True
- ○ False
5. Put these lines in order so the program counts a True Positive (inside the loop they are indented):
- 1. tp = tp + 1
- 2. for row in csv.DictReader(f):
- 3. if actual == "scab" and row["predicted"] == "scab":
- 4. actual = row["actual"]
Remember
Precision fears false alarms. Recall fears misses. Choose by the harm.