Class 10 · Unit 2: Modelling · Lesson 2.3 · 40 min
Classification, regression, clustering, association
Every samosa at break comes with a tea. Nobody told the canteen that — the receipts did.
Today you will: Predict a price with a trend line · Find groups nobody named · Discover what gets bought together
Story
Every day at break the school canteen runs out of tea before it runs out of anything else. The man who runs it, Gul Kaka, blames the weather.
Hiba has forty of his receipts from last week, with his permission. She reads them twice.
Hiba: “It isn't the weather. Look — every receipt with a samosa has a tea on it. Every single one.”
Bilal: “So?”
Hiba: “So on samosa days, make more tea.”
Bilal: “Did he tell you that?”
Hiba: “He didn't know it. The receipts did.”
Watch
Four kinds of model — — scan code 10.2.3 in the printed book.
Warm up your fingers
Skill: a formula with two conditions (5 min)
-
On typing.com: the symbols lessons.
-
Type three times, eyes on the screen:
=COUNTIFS(B2:B41;1;C2:C41;1) -
Goal this term: 38–44 WPM at 95%.
On the laptop
You need: LibreOffice Calc · datasets/mandi-prices.csv · datasets/apples.csv · datasets/canteen.csv · four-models.ods
Mission 1: Regression — predict a price (8 min, in pairs)
- ☐1
Open
mandi-prices.csv: the price of a box of apples at the mandi, for 16 weeks. - ☐2
Predict: what will a box cost in week 20? Write it in D1.
- ☐3
Make an XY (Scatter) chart. Right-click a point → Insert Trend Line → Linear → tick Show equation. Use the equation to work out week 20. How close was your guess?
Mission 2: Clustering — groups nobody named (8 min)
- ☐4
Open
apples.csv. Hide thevarietycolumn — now the data is unlabelled. - ☐5
Make an XY chart of weight against red_percent. Draw circles round the groups you see.
- ☐6
Unhide
variety. Did your groups match the three varieties?
Mission 3: Association — what goes with what (7 min)
- ☐7
Open
canteen.csv: 40 receipts; 1 means the item was bought. - ☐8
Count samosa receipts:
=COUNTIF(B2:B41;1). Count samosa and tea:=COUNTIFS(B2:B41;1;C2:C41;1). Now count tea on receipts without a samosa. - ☐9
Finish in H1: "Of the samosa receipts, … % had tea. Of the others, … %."
Mission 4: Name the model (2 min)
- ☐10
Sheet Problems in
four-models.ods: write classification, regression, clustering or association beside each of eight problems.
Finished early?
Find a second association in canteen.csv. Is it as strong as samosa and tea?
No laptop today?
Plot the 16 prices on squared paper and draw the best straight line by eye; read off week 20. Your teacher reads ten receipts aloud — tally samosa, and samosa-with-tea. Then name the eight problems aloud.
Now you know
- Classification predicts a class from a fixed list: Grade A, B or Reject. Discrete.
- Regression predicts a number that can take any value: a price, a temperature. Continuous.
- Clustering finds groups in unlabelled data that nobody named in advance.
- Association finds what goes with what: samosa → tea; bread → butter.
- Classification or clustering? Classification sorts into classes someone already named. Clustering discovers the groups itself.
Debate it
The canteen's receipts showed that samosa comes with tea. A big online shop does the same thing with everything you have ever bought.
- Is it the same? What is different when the receipts carry your name?
- Gul Kaka's receipts had no names on them. Why does that matter?
Check yourself — practice, not a test
1. Match each problem to its model:
Is this email spam or not? · What will a box of apples cost next week? · Which customers shop in a similar way? · People who buy bread also buy what?
Match with: Clustering · Classification · Regression · Association2. Predicting tomorrow's temperature in °C is:
- ○ Classification, because it is about weather
- ○ Regression, because the answer is a number that can take any value
- ○ Clustering, because temperatures group together
- ○ Association, because it depends on today
3. Clustering needs every example to be labelled first.
- ○ True
- ○ False
4. Classification works on ______ data; regression on continuous data.
5. Which are sub-categories of supervised learning? (choose all that apply)
- ○ Classification
- ○ Regression
- ○ Clustering
- ○ Association
Remember
A class, a number, a group, a pair — four questions, four kinds of model.