Class 6 · Unit 2: Data · Lesson 2.5 · 40 min
Good data, bad data
The weather app says: 29 °C, sunny, wear light clothes! Outside Srinagar it is snowing. The app worked perfectly — it just learned from Chennai.
Today you will: Hunt for missing and wrong data · Clean what can be cleaned · Write down where data came from
Story
A company builds an app to predict tomorrow's weather in Srinagar. They train it on weather data from Chennai.
In January the app says: "Tomorrow: 29°C, sunny. Wear light clothes!"
Outside, it is snowing. Everyone is wearing a pheran with a kangri underneath.
The app worked perfectly. The data was wrong for this place.
Watch
Why good data matters — — scan code 6.2.5 in the printed book.
Warm up your fingers
10 minutes on typing.com: bottom row continued.
- Continue the bottom-row lessons.
- Accuracy challenge: aim for 95%. Slow down if you need to. Careful typing prevents the "wrong data" problem you will see today.
On the laptop
Tools: LibreOffice Calc · class-survey-messy.ods (your class survey, "damaged" on purpose by your
teacher)
Mission 1: Find the problems (pairs, 8 min)
- ☐1
Open
class-survey-messy.odsfrom the shared folder. - ☐2
Predict: how many mistakes has your teacher hidden? Write your guess in H4.
- ☐3
Missing: find the empty cells and colour them yellow (Format → Cells → Background).
- ☐4
Wrong: find impossible values —
Sleep = 80,Crikcet,Season = Monsoon. Colour them red.
Mission 2: Fix, check and record (pairs, 10 min)
- ☐5
Fix what you can: change
CrikcettoCricket. An impossible number can't be guessed — delete it and leave the cell yellow. - ☐6
Data → AutoFilter on
Sleep: do the values look sensible now? - ☐7
Enough? Fair? In H1: is this enough rows to learn about all Class 6 students in Kashmir? In H2: if only walkers had answered, what would the travel chart wrongly say?
- ☐8
Record the source. In H3: who collected this data, when, and from whom. Save as
class-survey-clean.ods.
No laptop today?
Your teacher writes a 10-row messy table on the board. In pairs, circle the missing values, cross out the wrong ones, and write one sentence about what the whole table is missing.
Now you know
- Missing data: empty cells — the AI learns from less than you think.
- Wrong data: typing mistakes and impossible values teach the AI wrong things.
- Too little data: a few examples can't show the real pattern.
- Biased data: if some groups or places are missing, the AI is unfair or wrong for them.
- Cleaning fixes mistakes. Bias needs better collecting, not just cleaning.
- Ask where data came from. Data nobody can check shouldn't be trusted.
Checking the chain: a thousand-year-old method
Long before computers, scholars in the Muslim world faced a data problem that will sound familiar. Thousands of reports of the Prophet's ﷺ sayings were circulating. Some had been passed along carefully; others had been muddled, or invented. How do you tell which is which?
Their answer was a system called ʿilm al-ḥadīth, the science of ḥadīth. Its rules map almost exactly onto what you did today:
| What the scholars did | What a data scientist calls it |
|---|---|
| Isnād — every report had to carry its chain: who heard it from whom, all the way back | Provenance: where did this data come from? |
| Jarḥ wa taʿdīl — each narrator in the chain was investigated: memory, honesty, whether they ever contradicted themselves | Source checking: is this source reliable? |
| A report carried by many independent chains was stronger than one carried by a single person | Sample size: more sources, more confidence |
| Reports were graded ṣaḥīḥ (sound), ḥasan (good) or ḍaʿīf (weak) | Data quality rating |
Scholars such as al-Bukhārī (810–870) and Muslim ibn al-Ḥajjāj (c. 815–875) applied these rules to enormous collections of reports. Others compiled biographical dictionaries of narrators — thousands of entries recording who was careful, who was forgetful, and who could not be relied on. It was, in effect, a database of sources, built by hand.
So when you ask "Where did this number come from, and can I trust whoever wrote it down?" you are using a method your own region's scholarly tradition worked out more than a thousand years ago.
The grade ḥasan was formalised later, and is particularly associated with al-Tirmidhī (d. 892).
Debate it
An app that recognises apple diseases was trained only on photos of Delicious apples. A farmer in Sopore grows Ambri. What might happen, and who could be harmed?
Check yourself — practice, not a test
1. Match each problem to its example:
An empty cell where the sport should be · "Sleep: 80 hours" · Weather app trained only on Chennai data · Only 5 survey answers for all of Kashmir
Match with: Wrong data · Too little data · Biased data · Missing data2. An AI can be excellent even if its training data is full of mistakes.
- ○ True
- ○ False
3. The best way to fix biased data is usually to:
- ○ Delete the dataset
- ○ Colour the cells red
- ○ Collect more data from the groups that were left out
- ○ Make the chart bigger
4. "Garbage in, ______ out."
5. The chain showing who passed a report on to whom is called the ______. Today we would call it the data's source, or provenance.
6. A message about tomorrow's exam reaches you through four people, and nobody can say who started it. A ḥadīth scholar and a data scientist would both:
- ○ Share it quickly, before others do
- ○ Treat it as weak until the source can be checked
- ○ Believe it, because four people repeated it
- ○ Believe it only if it is written neatly
Remember
An AI is only as good as its data: check for missing, wrong, too little, and biased.