When data is presented in a grouped frequency table, individual values are unknown, so an exact mean cannot be found. Instead, use the midpoint of each class interval as a representative value, multiply each midpoint by its frequency, then divide the total by the overall frequency to find the estimated mean.

Why is the mean "estimated" for grouped data?

In a grouped frequency table, you know how many values fall within each class interval, but not the individual values. For example, if 8 people took between 10 and 20 minutes to complete a task, you do not know whether they took 11 minutes, 17 minutes, or anything in between. The best assumption — and the one GCSE expects — is that values are spread evenly within each class, making the midpoint the most representative single value for that class.

How do you find the midpoint of a class interval?

The midpoint of a class interval is the average of its lower and upper boundaries:

$$\text{midpoint} = \frac{\text{lower boundary} + \text{upper boundary}}{2}$$

Examples:

Class interval Midpoint calculation Midpoint
0 – 10 (0 + 10) ÷ 2 5
10 – 20 (10 + 20) ÷ 2 15
20 – 30 (20 + 30) ÷ 2 25
30 – 40 (30 + 40) ÷ 2 35

What is the step-by-step method?

  1. Find the midpoint of each class interval.
  2. Multiply each midpoint by its frequency to get f × x for each row.
  3. Find the total of the f × x column (this is Σfx).
  4. Find the total frequency Σf by adding all frequencies.
  5. Calculate: Estimated mean = Σfx ÷ Σf.

Full worked example: journey times

A survey recorded the journey times (in minutes) of 30 commuters.

Time (mins) Frequency (f) Midpoint (x) f × x
0 – 10 4 5 20
10 – 20 8 15 120
20 – 30 11 25 275
30 – 40 5 35 175
40 – 50 2 45 90
Total 30 680

Estimated mean = 680 ÷ 30 = 22.7 minutes (to 1 d.p.)

This means the average journey time is approximately 22.7 minutes, though the true mean could differ slightly because we used midpoints as approximations.

How do you handle open-ended or unequal class intervals?

Open-ended intervals (e.g. "50 or more") have no upper boundary, so a midpoint cannot be calculated in the usual way. If the question provides a suggested value, use it. If not, you may need to make a reasonable assumption — many GCSE questions avoid open-ended intervals for this reason.

Unequal class widths do not change the estimated mean method — still use the midpoint of each class regardless of width. However, they do affect how you draw a histogram (frequency density, not frequency, goes on the y-axis).

What mistakes should you avoid?

  • Using the class boundaries instead of the midpoint. Using 10 instead of 15 as the midpoint for the class "10–20" understates the contributions from that class.
  • Forgetting to sum the f × x column correctly. Set up the table neatly and double-check your arithmetic before dividing.
  • Dividing by the number of rows rather than the total frequency. The denominator is Σf (sum of all frequencies), not the count of rows in the table.
  • Reporting the exact mean. Always call it the estimated mean in your working and answer, to acknowledge that grouping has removed precision.

How does this appear in GCSE exam questions?

GCSE questions on this topic typically:

  • Give a complete grouped frequency table and ask for the estimated mean.
  • Sometimes ask you to "show that the estimated mean is approximately X" — set out all working clearly.
  • Occasionally give a partially completed table where you must fill in the midpoints or f × x values first.
  • May combine this with other statistics topics: for example, asking you to compare the estimated mean with the modal class or to suggest which class the median lies in.

Frequently asked questions

Is the estimated mean the same as the true mean?

Not necessarily. The estimated mean assumes values are concentrated at the midpoint of each class. If values cluster near one boundary — say, most commuters in the 10–20 minute class actually took close to 10 minutes — the true mean would be lower than the estimate. The estimated mean is the best answer available from grouped data alone.

What is the modal class?

The modal class is the class interval with the highest frequency — it is the group that contains the most values. It is not the same as the estimated mean. In the worked example above, the modal class is "20–30 minutes" (frequency 11), while the estimated mean is 22.7 minutes. These are different measures of average.

How do I find the class that contains the median from a grouped table?

Find the total frequency n and locate the class containing the (n ÷ 2)th value. For 30 values, the median is between the 15th and 16th value. Add cumulative frequencies: the first two classes hold 4 + 8 = 12 values, so the 15th and 16th values both fall in the "20–30" class. The median class is 20–30 minutes.

Can I use a calculator for the estimated mean in GCSE?

Yes — calculator papers at GCSE allow you to use a calculator. On a non-calculator paper, keep working exact (as fractions or in stages) until the final division, then convert. In the example above, 680 ÷ 30 = 22.6̄ recurring, rounded to 22.7 minutes (1 d.p.).


For Socratic statistics coaching — explaining why each step works, not just what to do — visit aitutors.me.