Data collection is the first stage of any statistical investigation. You must decide what to measure, how to record it accurately, and whether to gather your own data or use data already collected by someone else. Each choice affects how reliable your conclusions will be.

What is primary data?

Primary data is data you collect yourself, directly from the source.

Examples of primary data collection:

  • Carrying out a survey or questionnaire
  • Recording your own observations (e.g. counting cars passing a junction)
  • Conducting an experiment (e.g. measuring plant growth under different conditions)
  • Taking measurements (e.g. timing 100-metre sprints)

Advantages: You control exactly what is collected, in what format, and from whom. The data is fresh and specific to your question.

Disadvantages: Collecting primary data takes time and effort. Large surveys are expensive and difficult to organise. Response rates may be low.

What is secondary data?

Secondary data is data that someone else has already collected, which you use for your own analysis.

Examples of secondary data:

  • National census records
  • Weather data from the Met Office
  • Government statistics on population, health or employment
  • Data from a website or textbook

Advantages: Large, high-quality secondary datasets already exist and are quick to access. They may cover hundreds of thousands of people — far more than you could survey yourself.

Disadvantages: The data may not exactly match your question. You cannot control how it was collected, and it may be out of date.

What makes a good data collection sheet?

A data collection sheet (tally chart, observation sheet or questionnaire) should be:

Feature Why it matters
Clear headings Prevents confusion about what is being recorded
Pre-defined categories Speeds up recording and avoids inconsistent labels
A tally column Makes counting easy: five tallies = 1111 ̶1̶
A frequency column Summarises each category at a glance
A space for the total Allows a quick check that all data has been captured

Example data collection sheet for favourite subject:

Subject Tally Frequency
Maths
English
Science
Other
Total

Note the "Other" row — any response that does not fit the main categories still gets recorded.

How do you write unbiased questionnaire questions?

A good survey question is clear, unambiguous and does not push the respondent towards a particular answer.

Poor question (leading): "Don't you think our school canteen serves unhealthy food?"

  • This implies the food is unhealthy before the respondent answers.

Better question (neutral): "How would you rate the healthiness of our school canteen food?"

  • Very healthy [ ] Fairly healthy [ ] Neither [ ] Fairly unhealthy [ ] Very unhealthy

Rules for questionnaire design:

  1. One question per topic — do not ask two things at once ("Are you fit and healthy?").
  2. Give a suitable range of response options — do not overlap ("1–5" and "5–10") and always include a neutral option.
  3. Avoid technical language or jargon the respondent may not understand.
  4. Keep the question short — long questions are confusing.

What is bias in data collection?

Bias occurs when the data collected does not fairly represent the population being studied. Sources of bias include:

  • Sampling bias: only certain groups have the chance to respond (e.g. an online survey excludes people with no internet access).
  • Response bias: people answer differently from their true feelings (e.g. social desirability — saying they exercise more than they do).
  • Non-response bias: people who do not respond may be systematically different from those who do.
  • Observer effect: people change their behaviour when they know they are being observed.

What is a census versus a sample?

A census collects data from every member of the population. It gives complete, accurate information but is impractical for large populations (e.g. the UK National Census is held every ten years and takes enormous resources).

A sample collects data from a representative selection of the population. It is quicker and cheaper, but results carry uncertainty. The larger and more representative the sample, the more reliable the conclusions.

For most KS3 projects, a sample is appropriate. Aim for at least 30 data points; the larger the better.

Frequently asked questions

What is the difference between discrete and continuous data?

Discrete data can only take specific, separate values — usually counts. Number of siblings (0, 1, 2, 3, …) is discrete; you cannot have 2.7 siblings. Continuous data can take any value within a range. Height (1.73 m, 1.735 m, …) is continuous; there is always a possible value between any two measurements. The distinction affects which type of chart you use: bar charts for discrete data, histograms for grouped continuous data.

How large should my sample be?

There is no single rule, but a common guideline for school projects is a minimum of 30 responses, enough to produce meaningful averages and bar charts. For more reliable results, aim for at least 10% of the population (if the population is small) or a large absolute number if the population is very large. The key question is: does the sample include representatives from all important groups in the population?

Can I use data from the internet as secondary data?

Yes, but evaluate its reliability first. Ask: who collected it, when, how, and for what purpose? Government statistics and data from reputable research organisations are generally reliable. Data from unknown websites, opinion pieces or commercial surveys should be treated with caution. Always state your secondary data source in any statistics project.

What is a random sample and why is it better than a convenience sample?

In a random sample, every member of the population has an equal chance of being selected (e.g. putting names in a hat or using a random number generator). This minimises bias because no group is systematically favoured. A convenience sample selects whoever is easiest to reach (e.g. asking your classmates). It is quick but biased — your classmates are unlikely to be representative of the whole school, let alone the whole population.


For guided KS3 statistics help with Professor Pi, visit AI Tutors.