Short answer
A regular expression (regex) is a pattern of characters that describes a set of valid strings. For GCSE Computer Science, regex lets programs check whether user input — such as a postcode or email address — matches an expected format, making validation concise and precise.
At a glance
- Key stage
- GCSE
- Subject
- Computing
- Type
- Guide
- For
- Students
- Read time
- 5 min
- Last updated
- 8 October 2026
Where this fits
- Key Stage 3Years 7–9
- GCSEYears 10–11This article
Method at a glance
- ^ — start of string (no leading characters allowed)
- [A-Z]{1,2} — one or two uppercase letters (the area code, e.g
- \d{1,2} — one or two digits (the district, e.g
- \s — a single space separating the two halves
- \d — one digit (the sector)
- [A-Z]{2} — two uppercase letters (the unit)
- $ — end of string
What exactly is a regular expression?
Think of a regex as a template with wildcards. Rather than listing every valid postcode (SW1A 1AA, EC1A 1BB, …), you write one pattern that matches them all: a rule such as "two letters, one or two digits, a space, one digit, two letters." The computer then checks whether a given string fits the template.
In Python, the re module provides regex support. The two most common operations are:
re.match(pattern, string)— checks whether the pattern fits from the start of the string.re.search(pattern, string)— scans the whole string for any match.re.fullmatch(pattern, string)— requires the pattern to match the entire string.
What are the core building blocks of regex syntax?
| Symbol | Meaning | Example pattern | Matches |
|---|---|---|---|
. |
Any single character | c.t |
cat, cut, cot |
\d |
Any digit (0–9) | \d\d |
42, 07, 99 |
\w |
Any word character (letter, digit, underscore) | \w+ |
hello, x3, _val |
\s |
Any whitespace character | \s |
space, tab, newline |
[abc] |
Any one of the listed characters | [aeiou] |
a, e, i, o, u |
[a-z] |
Any character in the range | [A-Z] |
A, B, … Z |
[^abc] |
Any character NOT in the set | [^0-9] |
any non-digit |
^ |
Start of string | ^Hello |
string starting with Hello |
$ |
End of string | \.com$ |
string ending with .com |
What are quantifiers and how do they work?
Quantifiers specify how many times a character or group must appear.
| Quantifier | Meaning | Example | Matches |
|---|---|---|---|
* |
Zero or more | ab*c |
ac, abc, abbc |
+ |
One or more | \d+ |
1, 42, 9999 |
? |
Zero or one (optional) | colou?r |
color, colour |
{n} |
Exactly n times | \d{4} |
2026 |
{n,m} |
Between n and m times | [a-z]{2,5} |
ab, xyz, hello |
The quantifier applies to the character or group immediately before it. If you want a quantifier to cover multiple characters, use parentheses: (ab)+ matches ab, abab, ababab.
How do you write a regex for a UK postcode?
UK postcodes follow the format AB1 1CD (with variations). A simplified regex that covers most common cases is:
^[A-Z]{1,2}\d{1,2}\s\d[A-Z]{2}$
Breaking it down:
^— start of string (no leading characters allowed).[A-Z]{1,2}— one or two uppercase letters (the area code, e.g.SWorE).\d{1,2}— one or two digits (the district, e.g.1or12).\s— a single space separating the two halves.\d— one digit (the sector).[A-Z]{2}— two uppercase letters (the unit).$— end of string.
The pattern correctly matches SW1 1AA and E1 6AX but rejects SW11AA (missing space) and sw1 1aa (lowercase).
How are regular expressions used in programs?
Regular expressions are used whenever a program needs to validate, search, or extract structured text. Common school-level uses include:
- Input validation — reject a username before it reaches the database if it contains forbidden characters.
- Data extraction — pull all phone numbers from a block of text.
- Text replacement — swap every occurrence of an old domain name with a new one.
- Password strength checking — verify that a password contains at least one digit and one uppercase letter.
In Python, a simple validation function looks like this:
import re
def valid_email(address):
pattern = r'^[\w\.-]+@[\w\.-]+\.\w{2,}$'
return bool(re.fullmatch(pattern, address))
print(valid_email("student@school.ac.uk")) # True
print(valid_email("not-an-email")) # False
The r before the string prevents Python from misinterpreting the backslashes as escape sequences.
How do regular expressions compare with other validation approaches?
| Approach | Flexibility | Readability | Speed |
|---|---|---|---|
| Regular expressions | Very high | Low (to beginners) | Very fast |
if/in checks |
Low | High | Fast |
String methods (isdigit, etc.) |
Medium | High | Fast |
Dedicated library (e.g. email.utils) |
Task-specific | Medium | Fast |
Regex is most appropriate when the valid format has a clear, describable pattern. For highly complex rules (such as full RFC 5321 email validation), a dedicated library is usually clearer and more maintainable.
What common mistakes do students make with regex?
- Forgetting anchors — without
^and$, a pattern matches a substring:\d{4}matches the four-digit run insideabc1234xyz, which may not be the intention. - Forgetting to escape special characters — a literal dot in
.commust be written\.com; without the backslash,.matches any character. - Confusing
*and+—*allows zero matches, so\d*matches an empty string. Use+when at least one character is required.
Frequently asked questions
Is regex in the GCSE Computer Science specification?
Regular expressions appear in OCR's GCSE Computer Science specification (J277) as part of string handling and pattern-based validation. AQA's specification also references them in the context of input validation and string operations. You may be asked to interpret or write simple patterns in the examination.
Why do backslashes appear so often in regular expressions?
A backslash turns an ordinary character into a special code: \d means "any digit", \s means "any whitespace", and \n means a newline. Conversely, a backslash in front of a special regex character like . or * makes it literal again. In Python, placing an r before the string (a raw string) stops Python's own string-escape system from consuming the backslashes before the regex engine sees them.
Can you use regex in pseudocode for GCSE exams?
GCSE examinations rarely ask you to write full regex syntax in pseudocode. More commonly, you describe the validation rule in plain English or use a simplified notation. However, understanding what regex does helps you describe validation logic precisely and shows depth of knowledge.
What is the difference between re.match and re.fullmatch in Python?
re.match checks only whether the pattern fits at the beginning of the string — it succeeds even if characters follow the matched portion. re.fullmatch requires the pattern to cover the entire string. For validation, re.fullmatch (or anchoring with ^ and $) is usually correct; re.match is appropriate when you want to recognise strings that start with a particular format.
Want to practise writing and testing regular expressions with instant feedback? Professor Turing at aitutors.me can walk you through any pattern step by step.
Key terms
- Input validation
- Data extraction
- Text replacement
- Password strength checking
- Forgetting anchors
- Confusing * and +
- beginning
- entire