Detailed notes on Statistics and Probability for Cambridge Lower Secondary Mathematics, covering key concepts, explanations, examples, and exam-focused revision points.
Take this whole topic with you
Analysing Data — frequently asked questions
The things students keep getting wrong in this sub-topic, answered.
Analysing Data — Cambridge Lower Secondary Maths, Checkpoint
Revise the four big summary numbers — mean, median, mode and range — from lists, frequency tables and grouped tables, then use them to compare two sets of data with confidence ready for your Checkpoint.
At a glance
The mean is the total of all the values divided by how many there are.
The median is the middle value once the data is in order.
The mode is the value (or values) that appears most often.
The range is the largest value subtract the smallest.
From a frequency table, the mean uses ∑f∑fx — totals on both top and bottom.
From a grouped table you find the modal class and an estimated mean using midpoints.
Comparing two data sets needs one average and one spread — and a sentence in context.
A small range means more consistent data; a large range means more spread out.
What you’ll learn
Mapped to the Cambridge Lower Secondary Mathematics curriculum framework.
Find the mean, median, mode and range from a list of data.
Find these summaries from an ungrouped frequency table.
Find the modal class and an estimated mean from a grouped frequency table.
Compare two distributions using an average together with a spread.
The four summary numbers
Three averages and one spread — each one tells a different story about a data set.
When you have a list of numbers, four summaries do most of the heavy lifting.
The mean adds everything up and shares it out equally: mean=number of valuessum of values.
The median is the middle value once the list is in order. If there are two middle values, take the mean of those two.
The mode is the value that appears most often. A set can have no mode, one mode or more than one mode.
The range is a measure of spread: .
Summaries from an ungrouped frequency table
A frequency table puts repeated values in one row — adjust your method, not your formulae.
When data repeats, a frequency table is tidier than a list. The frequency f is just how many times that value x appears.
For the table below (shoe sizes of 20 students):
Shoe size x
4
5
6
7
8
Frequency
Grouped frequency tables and estimated mean
When data is grouped, you lose the exact values — so you estimate using midpoints.
If the data covers a wide range, it is often grouped into intervals. You can no longer read off exact values, so you find the modal class and an estimated mean.
Mass m (kg)
40≤m<50
50≤
Comparing two sets of data
Always compare with one average and one spread — and write the comparison in context.
When you compare two groups, the rule is simple: use one average (mean or median) and one spread (the range), and write a sentence that names the group and the units.
Suppose two classes sit the same 10-mark quiz:
Class A
Class B
Mean
6.4
7.1
Range
5
Choosing the right average
Each average has a sweet spot — and a weakness around outliers and category data.
Different averages work best in different situations.
Mean uses every value. It is the most informative when the data is fairly evenly spread, but a single huge or tiny value (an outlier) can pull the mean a long way.
Median is the middle value, so a single outlier barely moves it. The median is a safer choice for skewed data — like income or house prices.
Mode is the only average that works for categorical data (favourite colour, shoe brand, eye colour), because you cannot add up words.
For the salaries (in thousands) 20,22,24,25,110:
Mean = — dragged upwards by the outlier .
Where you'll use this next
Averages, range and comparisons feed straight into charts, probability and IGCSE statistics.
Mean, median, mode and range are the language of data — and they keep coming back:
Charts and graphs in the next subtopic let you see averages and ranges at a glance.
Probability uses relative frequency, which is built on the same idea of "how often does this value happen?".
Cumulative frequency and scatter graphs in later years extend the same comparison skills.
In real life, summaries are how journalists, scientists and economists describe huge data sets in a sentence.
Practise the four summaries on lists, tables and grouped tables until they feel automatic. The comparison sentence — average, spread, context — is the part that often makes the difference between a good and a top answer.
Bar and line charts visualise the same summaries.
Probability uses the language of relative frequency.
Comparisons reappear with cumulative frequency and box plots later.
Strong summary skills make every statistics topic easier.
Quick recap
Mean = total ÷ count; median = middle value; mode = most common.
Range = largest − smallest; small range means consistent data.
All resources on this platform are independently created by Tutopiya and have no endorsement from the International Baccalaureate Organization.
Where you'll use this next
range=largest−smallest
For the list 2,4,4,7,8:
Mean =52+4+4+7+8=525=5.
Median (middle of five values, sorted) =4.
Mode (most common) =4.
Range =8−2=6.
Three averages tell you where the data sits; the range tells you how spread out it is.
Each summary tells a slightly different story. Together they sketch the data quickly and accurately — which is exactly what you want when comparing two groups.
Mean = total divided by how many values there are.
Median = the middle value once the list is in order.
Mode = the most frequent value (a set can have none, one or several).
Range = largest value minus smallest value — a measure of spread.
f
2
5
7
4
2
The total frequency is ∑f=2+5+7+4+2=20.
Mode: the value with the highest frequency, so the modal shoe size is 6.
Median: with 20 values, the median is between the 10th and 11th values once they are in order. Adding the frequencies as you go — 2,7,14,… — both the 10th and 11th values fall in the size-6 row, so the median is 6.
Mean: multiply each value by its frequency, add the products, then divide by ∑f:
The trick is to remember that the frequency replaces writing the same value out multiple times. A working column for fx alongside the table keeps your numbers tidy and saves silly arithmetic slips.
Use xˉ=∑f∑fx for the mean from a frequency table.
The mode is the value with the highest frequency.
Find the median position using 2n+1, then count along the frequencies.
The range still uses the largest value minus the smallest.
m<
60
60≤m<70
70≤m<80
Frequency f
6
12
8
4
Modal class: the class with the highest frequency — 50≤m<60 (frequency 12).
Median class: with ∑f=30, the median is the 15th -ish value. Cumulative frequencies 6,18,26,30 tell you the 15th value is in 50≤m<60 — the median class.
Estimated mean: pick the midpoint of each class as x, multiply by the frequency, and divide:
xˉ≈∑f∑fx=3045(6)+55(12)+65(8)+75(4)=30270+660+520+300=301750≈58.3 kg.
The result is an estimate because every student in the 50≤m<60 class is treated as if they were exactly 55 kg. That is the trade-off when data is grouped — convenience now, a small loss of accuracy later.
The modal class is the interval with the tallest bar.
When you write the final answer, say it is an estimate — that single word reminds the reader that midpoints stand in for the real values inside each class.
Modal class = the class with the highest frequency.
Estimated mean uses midpoints: xˉ≈∑f∑fx.
Use cumulative frequency to identify the median class.
Grouped means are always estimates — say so when you write them.
9
A good comparison reads:
"Class B has a higher mean (7.1 marks vs 6.4 marks), so on average they scored higher than Class A. However, Class A has a smaller range (5 vs 9), so their scores were more consistent."
Two ideas pop out. The average answers "who did better on the whole?" — a higher mean (or higher median) means higher scores on average. The range answers "who was more consistent?" — a smaller range means the scores cluster more tightly.
Class B scored higher on average but was more spread out than Class A.
Two reminders to keep your comparison full marks. Always use the same pair of statistics for both groups (do not compare A's mean with B's median). And always write a sentence in context — name the groups, give the numbers and the units, and say what the difference means in plain English.
Compare one average AND one spread — never just one.
Higher mean / median = higher on average.
Smaller range = more consistent values.
Write the comparison in context, with units and group names.
40.2
110
Median =24 — the middle value, unaffected by the outlier.
The median is the better summary of "typical" salary here. The mean tells you about the total payroll, but the median tells you about the middle person. Picking the right average is part of analysing data — and showing you understand the difference is exactly what comparison questions reward.
Mean uses every value — great when data is balanced, sensitive to outliers.
Median is the middle value — robust to outliers, good for skewed data.
Mode is the only average for categorical (non-numeric) data.
Always think: which average best represents typical?
From a grouped table find the modal class and the estimated mean using midpoints.
Compare two groups with one average AND one spread, in context.
Outliers move the mean a lot; the median barely shifts.
Mode is the only average that works for categorical data.
8
Step 3
With seven values, the middle one is the 4th — that is 6, so the median is 6.
Step 4
Add up all seven values, then divide by 7 for the mean.
xˉ=73+4+5+6+8+8+8=742=6
Step 5
Range is largest minus smallest.
8−3=5
Step 1
Sort the list into order from smallest to largest.
7,9,10,12,14,15
Step 2
There are six values, so the median is between the 3rd and 4th.
Step 3
Take the mean of 10 and 12.
210+12=11
Answer
Median =11.
f
4
9
7
3
2
Work out the mean number of pets per student.
Step-by-step solution
Step 1
Check the total frequency.
∑f=4+9+7+3+2=25
Step 2
Calculate fx for each column, then add them up.
∑fx=0(4)+1(9)+2(7)+
Step 3
Use the mean formula for a frequency table.
xˉ=∑f∑fx
Answer
Mean =1.6 pets per student.
<
10
10≤t<20
20≤t<30
30≤t<40
Frequency (f)
8
14
12
6
Estimate the mean time.
Step-by-step solution
Step 1
Find the midpoint of each class — half-way between the lower and upper boundaries.
x=5,15,25,35
Step 2
Total frequency.
∑f=8+14+12+6=40
Step 3
Calculate ∑fx using the midpoints.
∑fx=5(8)+
Step 4
Divide to estimate the mean.
xˉ≈40760=19
Answer
Estimated mean ≈19 minutes.
=
4
Class Q: mean =12.5, range =9.
Write a comparison of the two classes using the mean and the range.
Step-by-step solution
Step 1
Compare the means first — Class Q has the higher mean, so on average they scored higher.
Step 2
Then compare the ranges — Class P has the smaller range, so their scores were more consistent.
Step 3
Write the comparison in a clear sentence with the numbers and units.
Answer
On average Class Q scored higher than Class P (mean 12.5 vs 11.2 marks). However, Class P's scores were more consistent because their range (4 marks) was smaller than Class Q's (9 marks).
x
x
60
3
(
3
)
+
4(2)=
0+
9+
14+
9+
8=
40
=
2540=
1.6
15
(
14
)
+
25(12)+
35(6)=
40+
210+
300+
210=
760
Analysing Data — Cambridge Lower Secondary Mathematics — Checkpoint Revision Notes & Practice | Tutopiya