Class 8 · Unit 2: Building with no-code AI · Lesson 2.4 · 40 min
How good is it, really?
“Ours is 96%!” “On what?” One bottle, one moment. On fifteen held-back items: 80% — and it missed every battery.
Today you will: Score a model on unseen items · Break accuracy down by class · Decide which mistake your project can't afford
Story
Bilal: “Ours is ninety-six per cent.” (before the teacher has even asked.)
Hiba: “On what?”
Bilal looks at his screen. Ninety-six is the number Teachable Machine showed when he held up a plastic bottle. One bottle. One moment.
So they do it properly. Fifteen held-back items, one at a time, writing down what the model said and what the item actually was. Twelve right, three wrong. Not ninety-six per cent — eighty.
And all three mistakes are the same mistake: a hazardous item called dry waste — three of the five hazardous items.
Bilal: “Eighty sounds fine, until you notice it missed every battery in the test.” (slowly)
Watch
Measuring a model honestly — — scan code 8.2.4 in the printed book.
Warm up your fingers
Skill: numbers and symbols, Term 1 (5 min)
- On typing.com, continue this term's lessons, then practise the
row you reach for least:
= ( ) / * % - Type this line three times without looking down:
=COUNTIF(C2:C16;"yes")/15*100 - Target: 28–32 WPM at 93%+. Accuracy first — a mistyped formula costs more than a slow one.
On the laptop
You need: your model and held-back items from Lesson 2.3 · model-scores.ods
Mission 1: Count (10 min, same groups)
- ☐1
Columns: Item · What it really is · What the model said · Right?
- ☐2
Show each held-back item once. Type what it said. Don't retake a photo until it agrees — that's cheating your own test.
- ☐3
Predict your accuracy. Then: correct
=COUNTIF(D2:D16;"yes")in F2, accuracy=F2/15*100in F3.
Mission 2: Look at every mistake (10 min)
- ☐4
Sort by Right?. For each mistake: Confused with · Why you think so — light, background, angle, too few examples.
- ☐5
Accuracy per class:
=COUNTIFS(B2:B16;"hazardous";D2:D16;"yes")/COUNTIF(B2:B16;"hazardous")*100— and the other two. - ☐6
Finish: "Our model is worst at ___, because ___." Mark any mistake that would be dangerous in real use costly, in red.
No laptop today?
The teacher reads 15 printed results. Tally right and wrong, work out the percentage, then split the tally by class. Two groups both scored 80% — would you trust both with the school's battery bin?
Now you know
- Accuracy = correct ÷ total × 100, on data the model has never seen — that's model evaluation.
- Confidence is not accuracy. One prediction's preference isn't a record of many.
- One number hides things. Always break accuracy down by class.
- Mistakes are not equal. Decide which one you can't afford before you claim it works.
- Then refine: add data, remove wrong data, change settings, retrain — and test on the same set.
Debate it
Two groups build a hazardous-waste sorter. Group A is 90% accurate overall but catches only half the hazardous items. Group B is 80% overall and catches every hazardous item.
Which model should the school use? Who should decide — the group that built it, the teacher, or the people who empty the bins?
Check yourself — practice, not a test
1. A model shows 96% beside one prediction. This number is:
- ○ The model's accuracy
- ○ That one prediction's confidence score
- ○ The percentage of training photos it remembers
- ○ The number of classes it can tell apart
2. Accuracy is worked out as correct ÷ ______ × 100.
3. A model that scores 80% overall must be at least roughly 80% on each of its classes.
- ○ True
- ○ False
4. Which make a test honest? (choose all that apply)
- ○ Using items the model never trained on
- ○ Retaking the photo until the model agrees
- ○ Writing down the wrong answers as well as the right ones
- ○ Working out the accuracy for each class separately
5. Put model evaluation in order:
- 1. Look at each mistake and decide what to fix
- 2. Break the score down by class
- 3. Write down what it said and what the item really was
- 4. Show the model an item it has never seen
- 5. Count the correct answers and divide by the total
Remember
"How accurate is it?" is half a question. The other half is: on what data, and which class does it fail?