Class 10 · Unit 4: Computer vision · Lesson 4.5 · 40 min
Inside a CNN
Take your feature map from last lesson. Two formulas later you'll have built three layers of a real neural network.
Today you will: Apply ReLU to a feature map · Shrink it with max pooling · Trace an image through all four layers of a CNN
Story
Bilal has the apple's feature map up on his screen, all reds and greens.
Bilal: “So is this it? Is this how the phone knows it's an apple?”
Hiba: “It's the start. That's just the edges. A real network does that with dozens of kernels, then throws half the numbers away, then shrinks what's left, then does it all again.”
Bilal: “Throws numbers away? On purpose?”
Hiba: “On purpose. It keeps what matters. Starting with those.” (She points at the red band down the left side.)
Watch
Four layers, one apple — — scan code 10.4.5 in the printed book.
Warm up your fingers
Skill: a formula inside a formula (5 min)
-
On typing.com: the symbols lessons.
-
Type three times, eyes on the screen:
=MAX(OFFSET($ReLU.$A$1;(ROW()-1)*2;(COLUMN()-1)*2;2;2)) -
Count the brackets: every one you open, you close. Goal this term: 38–44 WPM at 95%.
On the laptop
You need: LibreOffice Calc · your convolution.ods from Lesson 4.4
Mission 1: ReLU (8 min, in pairs)
- ☐1
Add a sheet named ReLU. In A1:
=MAX(0;$Output.A1). Fill it across to J and down to 10. - ☐2
Predict: what happens to the red band down the left of the apple? Write it down.
- ☐3
Give ReLU the same colour scale as Output. Were you right?
Mission 2: Max pooling (10 min)
- ☐4
Add a sheet named Pool. In A1, paste the formula from the warm-up. It finds the biggest number in each 2 × 2 block of ReLU. Fill it across to E and down to 5.
- ☐5
Predict: the pooled map is 5 × 5. Will you still be able to see the edge of the apple?
- ☐6
Colour it. Count: how many numbers did you start with, after convolution, and after pooling?
Mission 3: Fully connected, on paper (5 min)
- ☐7
Read your 5 × 5 pool row by row into one list of 25 numbers — that is flattening.
- ☐8
A fully connected layer gives every one of the 25 numbers a weight for "apple" and a weight for "not apple", and adds them up. Which positions in your list would you give a big "apple" weight to? Why?
Finished early?
Change the pooling to average pooling: replace MAX with AVERAGE.
Compare the two 5 × 5 maps. Which keeps the edge sharper?
No laptop today?
Your teacher writes a 4 × 4 feature map with some negative numbers on the board. The class does ReLU aloud (every negative → 0), then max pooling into a 2 × 2, one block per row of desks. Then flatten the 2 × 2 into a list of 4 and discuss which should count most towards "apple".
Now you know
- A CNN is a deep-learning model for images, built from four kinds of layer.
- Convolution: many kernels make many feature maps — edges first, then shapes, then objects.
- ReLU: every negative becomes 0; positives stay. The useful features stand out.
- Pooling: shrink the map, keep what matters. Max pooling keeps the biggest in each patch.
- Fully connected: flatten to one list, weigh every number, give a score for each label.
- Nobody chooses the kernels. The network learns them from labelled images, by adjusting its weights after every mistake.
Debate it
A CNN that recognises faces learned its kernels from millions of photos. Nobody can point to the kernel that means "eyes".
- If a face-recognition system wrongly matched a person to a crime, could anyone explain why? (Lesson 2.4's Think asked the same about neural networks.)
- Should a system nobody can explain be used to make decisions about people? Where would you draw the line?
Check yourself — practice, not a test
1. Put the layers of a CNN in order:
- 1. Convolution
- 2. Pooling
- 3. Fully connected
- 4. ReLU
2. What does ReLU do to a feature map?
- ○ Makes it smaller
- ○ Turns negative numbers into 0 and keeps the positives
- ○ Adds a kernel
- ○ Gives the final label
3. Max pooling on the 2 × 2 block (3, 7, 0, 5) gives:
- ○ 3.75
- ○ 0
- ○ 7
- ○ 15
4. In a CNN, a person chooses every kernel by hand.
- ○ True
- ○ False
5. The output of a convolution layer is called a feature ______.
Remember
Find features, drop the negatives, shrink, then vote.