Class 9 · Unit 3: Maths for AI · Lesson 3.4 · 40 min
Probability inside an AI decision
“Scab — 98% confident.” “Scab — 52% confident.” Same word. Should the uncle spray both trees?
Today you will: Read a confidence score as a probability · Find where the mistakes cluster · Set a line for when a person checks
Story
The leaf app gives Bilal's uncle two answers in a row.
Scab — 98% confident. He sprays that tree.
Scab — 52% confident. He is about to spray that one too, when Hiba stops him.
Hiba: “That second one is not the app telling you it is scab. That is the app telling you it could not decide.”
Both answers said the same word. Why should he treat them differently?
Watch
The number beside the answer — — scan code 9.3.4 in the printed book.
Warm up your fingers
4 minutes on typing.com: accuracy under thought.
- Term 2 lessons. Target 94%.
- Then type
COUNTIFSand>=ten times each. Today's formulas use both, and>=typed as=>is the error you will hunt for five minutes.
On the laptop
Tool: LibreOffice Calc. File needed: unit-3/leaf-predictions.ods — twenty
predictions from the leaf classifier the class trained in Class 8. Column B is what the model
said, column C how confident it was, column D what the leaf actually was.
Mission 1: Your mission
- ☐1
In E2 type
=IF(B2=D2;"Right";"Wrong")and fill down to E21. - ☐2
In G2, overall accuracy:
=COUNTIF(E2:E21;"Right")/20*100. Write it down. - ☐3
Now split by confidence. In G4 count the high-confidence predictions:
=COUNTIFS(C2:C21;">=90"). In G5, how many of those were wrong:=COUNTIFS(C2:C21;">=90";E2:E21;"Wrong"). - ☐4
Repeat for the middle band in G7 and G8, using
">=70"together with"<90". - ☐5
Repeat for the low band in G10 and G11, using
"<70". - ☐6
Fill in this sentence in G13: "The model was wrong ____ times in twenty. ____ of those mistakes were predictions it was less than 70% sure of."
- ☐7
Make the rule. In G15, count how many photos fall below 70%:
=COUNTIFS(C2:C21;"<70"). In G16 write what fraction of the work that is, as a percentage. - ☐8
In G17 answer: if a person checks only those photos, how many of the model's mistakes would they catch, and how much of the work is that?
- ☐9
Try a stricter rule. Change the line from 70 to 90 and answer the same two questions. Which rule would you give the uncle, and why?
No laptop today?
Print the twenty predictions on cards and lay them on a desk in confidence order, highest to lowest. Turn over each card to reveal whether it was right. The wrong ones cluster at one end of the row, and the class can see the line where a person should start looking.
Now you know
- A confidence score is a probability, not a verdict. "Scab, 98%" and "Scab, 52%" are different answers.
- The scores add to 100% the model must split its belief between the classes it knows. It can't say "I can't tell".
- Mistakes cluster at low confidence. The score tells you where to doubt.
- A threshold turns a probability into a plan: "below 70%, a person looks".
- Stricter lines catch more mistakes and cost more work. Where to draw the line is a human decision, not a mathematical one.
🕌 From our heritage
Long before machines printed percentages, Muslim jurists were careful to distinguish yaqīn — certainty — from ẓann, probable knowledge, and from mere doubt. They held that much practical reasoning rests on ẓann al-ghālib, what is most probably the case, and they said so openly rather than dressing a likely answer as a certain one. A model's 52% is ẓann. The honesty is in naming it.
A standard distinction in uṣūl al-fiqh (the principles of jurisprudence); see the discussion of certainty and probability in the classical uṣūl literature. You met the branching questions of that tradition in Class 6.
Debate it
A hospital screening tool flags possible disease in chest X-rays. It is 94% accurate overall, and its mistakes cluster below 75% confidence.
Where would you set the line for a doctor to look, and what would it cost to set it too high? Answer for a hospital with two radiologists and four hundred X-rays a day.
Check yourself — practice, not a test
1. A model outputs "Scab 52%, Healthy 48%". The best reading is:
- ○ The leaf is definitely scabbed
- ○ The model is 52% accurate
- ○ The model could not decide between the two
- ○ The model is broken
2. In a two-class model, the confidence scores always add up to ______%.
3. A model's mistakes are spread evenly across all confidence levels.
- ○ True
- ○ False
4. Moving the human-check line from 70% up to 90% will:
- ○ Catch fewer mistakes and cost less work
- ○ Catch more mistakes and cost more work
- ○ Catch more mistakes and cost less work
- ○ Make no difference to either
5. Why is a confidence score useful? (choose all that apply)
- ○ It shows where the model is least reliable
- ○ It lets a team decide what a human should check
- ○ It proves the model is right
- ○ It turns one answer into a probability you can reason about
Remember
A confidence score tells you where to doubt. The line for a human to look is your decision, not the model's.