Class 8 · Unit 4: Data and fairness · Lesson 4.1 · 40 min
Where AI gets its data
2.4 million plant photos trained the app. Nobody says who took them. Somewhere in there might be your chinar.
Today you will: Write a dataset card · Spot the questions a dataset can't answer · Decide which data you'd trust
Story
Bilal finds a free app that promises to name any plant from a photo. It is fast, and it is right most of the time.
Hiba: “Where did it learn all that?”
Bilal shrugs. "From the internet?"
They read the app's page properly. It says the model was trained on 2.4 million plant photographs. Nowhere does it say who took them, which countries they came from, or whether the photographers were asked.
Hiba thinks about the photo of their own chinar tree that Bilal posted last spring.
Hiba: “So somewhere in those millions, there might be ours.”
Watch
Every dataset has a history — — scan code 8.4.1 in the printed book.
Warm up your fingers
Skill: typing while you think (Term 2 focus, 6 min)
- On typing.com, do one Intermediate paragraph lesson.
- Then, in Writer, type these four questions from memory, as fast as you can while staying accurate: Who collected it? When? Who is in it? Did they agree?
On the laptop
You need: LibreOffice Calc · dataset-card.ods (nine questions) · mystery-datasets.odt
Mission 1: Your mission
- ☐1
Save
dataset-card.odswith your names. Card 1: fill it for a dataset your class made — the leaf photos, the waste photos. Nine questions: What's in it? · Who collected it, when? · Where? · Who or what is in it? · Were they asked? · Who labelled it? · Who's missing? · What was it made for? · Who benefits? - ☐2
Predict: how many not stated will the two mystery datasets have?
- ☐3
Cards 2 and 3: fill one card each for the mystery datasets. Where the description is silent, write not stated — don't guess.
- ☐4
Count the not stated boxes across all three cards.
- ☐5
One sentence: which dataset would you trust for a decision that matters — and why?
No laptop today?
Your teacher reads the three descriptions aloud. Rule the card into your notebook, fill it, and write not stated where the description is silent. Count the blanks together.
Now you know
- Data has a history. Someone chose what to collect, whom to ask and what to leave out — and it ends up inside the model.
- Data arrives three ways: collected on purpose, found (scraped from the web), or left behind by people using a service.
- Labels are human judgements. Different people label differently.
- Consent means the people agreed, knowing the use. Posting a photo isn't agreeing to train a model.
- Class 6's isnād question, new subject: who does this come from, and how do we know?
Debate it
The plant app's 2.4 million photos help farmers identify diseases early. Nobody who posted those photos was asked.
Does a good result make it acceptable? What would a fair way to build that dataset have looked like?
Check yourself — practice, not a test
1. What does a dataset card record?
- ○ How fast a model runs
- ○ Where the data came from, who is in it, who labelled it and what is missing
- ○ The price of the software
- ○ How many people use the app
2. Which of these are ways data reaches a model? (choose all that apply)
- ○ Collected on purpose by someone measuring or photographing
- ○ Found on the web and taken without asking
- ○ Left behind as a by-product of people using an app
- ○ Generated by the laptop when it is switched off
3. If a photo is public on the internet, the person who posted it has agreed to it being used to train an AI model.
- ○ True
- ○ False
4. Deciding that one leaf counts as "scab" and another as "healthy" is a human ______, not a fact of nature.
5. On a dataset card, writing "not stated" means:
- ○ The dataset is illegal
- ○ We do not know, and that gap is worth recording
- ○ The answer is zero
- ○ The data is fake
Remember
No dataset falls from the sky. Every row was a choice somebody made.