Regular expressions for text work
A regex is a search pattern that matches shapes of text instead of exact text — the everyday power tool for finding, masking and restructuring strings.
- 7 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A regular expression is a search pattern that describes the shape of text.
Think "an @ with words on both sides", rather than the exact text itself.
Think of a cookie cutter. You do not find star-shaped cookies by comparing dough to one specific cookie. You press a star-shaped cutter anywhere in the dough, and everything matching the shape comes out. A regex (short for regular expression) is a cookie cutter for strings.
Why it exists
Ordinary search finds exact text: searching "priya@example.com" finds that one address. But cleanup jobs need every email address, every phone number, every date — an unlimited family of strings sharing one shape.
Describing shapes needs a small language. In it, symbols stand for character families. One symbol means "any digit", another means "any letter". Other markers mean "one or more of these" and "at this position". A few characters of pattern can match millions of possible strings.
How it works
The pattern reads left to right, like the text it hunts:
pattern: [\w.]+ @ [\w.]+
letters, an @ letters,
dots (1+) dots (1+)
text: "mail priya@example.com today"
└────── match ──────┘Each piece of the pattern consumes a stretch of text. If every piece finds its stretch in order, the whole pattern matches at that spot. The search engine slides the cutter along the string and stamps out every match.
A real example you have seen
When a chat app turns a typed link blue and clickable, a URL-shaped pattern found it. When a form instantly says "that does not look like a phone number", a pattern failed to match. When your email app hides your card number in a receipt as "XXXX 4321", a digit-shaped pattern found and masked it.
Remember this
- A regex matches the shape of text, not exact text.
- Symbols stand for families: any digit, any letter, repeated, at the start.
- Regex is for shapes; it cannot understand meaning — that is a model's job.
What to learn next
- spaCy pipelines — where rule-based matching hands over to linguistic models.
- Text normalisation — the cleanup passes regex substitutions implement.
- Named entity recognition — finding people and places when shape alone is not enough.
Developer — Code and libraries.
Setup
No installs. re ships with Python. Verified on Python 3.10.
Find, extract, mask — the three daily jobs
import re
log = """user=priya@example.com status=FAILED at 2026-08-30 14:02:11
user=arjun@shop.co.in status=OK at 2026-08-30 14:05:37
user=bad-entry status=FAILED at 2026-08-30 14:06:02"""
# findall: every email-shaped string
emails = re.findall(r"[\w.+-]+@[\w-]+\.[\w.]+", log)
print("emails:", emails)
# named groups turn a line into a dictionary
pattern = re.compile(
r"user=(?P<user>\S+) status=(?P<status>\w+) at (?P<time>[\d: -]+)")
for m in pattern.finditer(log):
if m["status"] == "FAILED":
print("failed login:", m["user"], "at", m["time"])
# sub: clean text instead of finding it
masked = re.sub(r"[\w.+-]+@[\w-]+\.[\w.]+", "<EMAIL>", log.splitlines()[0])
print("masked:", masked)emails: ['priya@example.com', 'arjun@shop.co.in'] failed login: priya@example.com at 2026-08-30 14:02:11 failed login: bad-entry at 2026-08-30 14:06:02 masked: user=<EMAIL> status=FAILED at 2026-08-30 14:02:11
The walkthrough
The email pattern, piece by piece. [\w.+-]+ means "one or more word characters, dots, pluses or hyphens" — the part before the @. Then a literal @. Then [\w-]+\.[\w.]+: domain, a literal dot (escaped as \. because a bare dot means "any character"), then the ending. Good enough for log cleanup; deliberately not a full validator.
The r"..." prefix is not optional style. It marks a raw string, stopping Python from eating backslashes before the regex engine sees them. Without it, "\w" may work today and "\d" break tomorrow. Every regex, always, gets the r.
Named groups (?P<user>...) beat numbered ones. m["user"] still reads correctly six months later; m.group(1) breaks the day someone adds a group at the front.
Second output line is a lesson in disguise. "bad-entry" is not an email, but \S+ — "one or more non-space characters" — matched it happily. The pattern found what the log format promised, not what you assumed lived there. Patterns are contracts with the text's shape; wrong assumptions surface as wrong matches, silently.
Common mistakes
Greedy matching eating too much. + and * grab the longest possible stretch. On user=a status=b user=c status=d, the pattern user=(.*) status= captures a status=b user=c. Add ? for the shortest stretch: (.*?), or better, match \S+ instead of .*.
Parsing HTML with regex. Nested tags are beyond what regular patterns can express. For a quick scrape of one known page, fine; for anything real, use an HTML parser (selectolax, BeautifulSoup) and save regex for the text inside the tags.
Validating emails "properly" with regex. The real email standard allows horrors no sane pattern covers. Match the rough shape, then verify by sending a confirmation — every serious system does exactly this.
Rewriting instead of composing. A 200-character regex nobody can read should have been three small patterns and an if. Complexity you cannot read at review time is complexity you cannot debug at 2 a.m.
Try it yourself
Extract every timestamp from the log with one findall. Then write a pattern for Indian mobile numbers — ten digits, optional +91 prefix, optional spaces — and test it against "+91 98765 43210", "9876543210" and the trap "12345".
What to learn next
- spaCy pipelines — where rule-based matching hands over to linguistic models.
- Text normalisation — the cleanup passes regex substitutions implement.
- Named entity recognition — finding people and places when shape alone is not enough.
Researcher — Mathematics and papers.
The theory underneath
Classical regular expressions describe exactly the regular languages — Kleene (1956); every pattern compiles to a nondeterministic finite automaton (Thompson, 1968) and can run in O(n·m) time for text length n and pattern size m, or O(n) after DFA construction. Practical engines diverge from theory in both directions. Backtracking engines (Python's re, PCRE, Java) support features beyond regular languages — backreferences make matching NP-hard in general (Aho, 1990) — and pay for it with exponential worst cases: nested quantifiers like (a+)+$ on "aaaa...b" exhibit catastrophic backtracking, the basis of ReDoS denial-of-service attacks (OWASP tracks these; Cloudflare's 2019 outage was one). Automata-based engines (RE2 — Cox, 2010; Rust's regex; Hyperscan) guarantee linear time by dropping backreferences. Python offers no built-in linear-time engine; the third-party google-re2 binding provides one, and Python 3.11+ added atomic groups (?>...) and possessive quantifiers ++ for manual mitigation.
Unicode semantics
re operates on code points with approximate Unicode classes: \w covers letters/digits/underscore across scripts by default on str. It lacks full UTS #18 support — no \p{Script=Devanagari} property syntax and no grapheme-cluster awareness (an emoji family sequence is several code points; . matches one). The third-party regex module supplies \p{...}, \X (grapheme match), and approximate (fuzzy) matching. Case-insensitive matching under full Unicode case folding (Turkish ı, German ß) also behaves better there. For pipelines feeding models, remember regex runs after decoding and before tokenization — GPT-family tokenizers themselves use a large regex as the pre-tokenization split rule (visible in tiktoken's source), so this material lives inside modern LLM stacks, not beside them.
Regex in ML preprocessing practice
Standard roles: PII masking (emails, phones, IDs) before training or logging; template/boilerplate stripping in web corpora (the C4 cleaning rules — Raffel et al., 2020 — are largely regex filters: line-level punctuation checks, "lorem ipsum" and brace filters); weak supervision, where patterns act as labelling functions (Ratner et al., 2017, Snorkel) whose noisy votes train a model that generalises past the patterns; and tokenizer pre-splitting as above. Engineering guidance that survives scale: precompile with re.compile in hot loops, prefer anchored and literal-prefixed patterns (engines optimise literal scanning), cap quantifier nesting, and fuzz any user-facing pattern with adversarial inputs for ReDoS before deployment.
What to learn next
- spaCy pipelines — where rule-based matching hands over to linguistic models.
- Text normalisation — the cleanup passes regex substitutions implement.
- Named entity recognition — finding people and places when shape alone is not enough.