Class 10 · Unit 3: Evaluating models · Lesson 3.5 · 40 min
Bias, transparency and accountability
83% accurate — for everyone? Split the results by orchard and a different model appears.
Today you will: Evaluate a model group by group · Write a model card · Say who answers when a model is wrong
Story
The leaf app's makers are pleased. 83% on sixty new photos, tested properly this time.
Hiba's uncle is not. "It's worse up here," he says. His orchard is on the hill, where the trees are in shade half the day. "Down in the valley my cousin says it's very good. Up here it gets it wrong all the time."
Bilal: “The sign says eighty-three.”
Uncle: “For whose trees?”
Hiba looks at the file. Every photo has a column she hasn't used yet: orchard — Valley or Hill.
Hiba: “Nobody ever split it.”
Watch
Honest about who it works for — — scan code 10.3.5 in the printed book.
Warm up your fingers
Skill: a count with three conditions (5 min)
-
On typing.com: the symbols lessons.
-
Type three times, eyes on the screen:
=COUNTIFS(B2:B61;"Hill";C2:C61;"scab";D2:D61;"scab") -
Goal this term: 38–44 WPM at 95%.
On the laptop
You need: LibreOffice Calc · LibreOffice Writer · datasets/leaf-predictions.csv · model-card.odt
Mission 1: Split it (12 min, in pairs)
- ☐1
Open
leaf-predictions.csv. Column B says which orchard each photo came from. - ☐2
Predict: will the model be more accurate on Valley or Hill photos — and by how much? Write it down.
- ☐3
For Valley only, build the confusion matrix with
COUNTIFSand three conditions, like the warm-up. Work out accuracy and recall. - ☐4
Do the same for Hill. Put the two side by side with the overall results.
Mission 2: The model card (13 min)
- ☐5
Open
model-card.odt. Fill each heading in one or two lines:- What it does — and what it must not be used for
- What it learned from — say where the photos came from
- How it was tested — how many photos, which orchards
- Results — overall and by orchard: accuracy and recall
- Known weaknesses
- Who is accountable — and how a farmer reports a wrong answer
- ☐6
Under Known weaknesses, write the sentence a hill farmer most needs to read.
Finished early?
Name two other groups this model might fail for, that the file can't tell you about. How would you collect the data to check?
No laptop today?
Your teacher writes the two small matrices on the board (Valley and Hill). Pairs work out accuracy for each, then write the model card on paper, one heading per pair, and read the headings aloud in order.
Now you know
- Bias: a model that works better for some groups than for others. An overall score can hide it completely.
- Evaluate group by group. Split the test results by any group that could matter — place, season, language, device — and compare.
- Transparency: people affected by a model can see how it was tested and how well it works for people like them.
- Accountability: someone named answers for its mistakes, and there is a way to report one.
- A model card puts all of this on one page that travels with the model.
Debate it
The makers could fix the hill problem by collecting shade photos from hill orchards and retraining.
- Until then, should the app be used on the hill at all? Who should decide?
- The sign still says 83% accurate. What should it say now?
Check yourself — practice, not a test
1. An overall accuracy of 83% can hide:
- ○ Nothing, if the test was done on unseen data
- ○ A group the model fails
- ○ The number of training photos
- ○ The positive class
2. Match each ethical concern to what it asks:
Does it work equally well for every group? · Can people see how it was tested? · Who answers for its mistakes?
Match with: Accountability · Bias · Transparency3. To find hidden bias, you can split the test results by group and compare.
- ○ True
- ○ False
4. A one-page report that travels with a model, giving its purpose, data, results and weaknesses, is called a model ______.
5. Why did the leaf model do worse on hill photos?
- ○ Hill leaves never get scab
- ○ The hill photos were taken at night
- ○ Most of its training photos came from sunny valley orchards
- ○ The model was not tested
Remember
A score for everyone can hide a failure for someone. Split it.