When users write 'f r e e m0ney' to dodge your filter
People rewrite banned words with spaces, symbols and lookalike characters specifically to slip past a keyword filter, and evasion-aware normalisation catches it back.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Filter evasion is writing a banned word in disguise — "f r e e m0ney" instead of "free money." It is written that way so a keyword filter does not recognise it.
Think about passing a note in a classroom where certain words get you in trouble. You do not write the forbidden word plainly. You write it in a way a quick glance would miss. A friend reading closely still understands perfectly.
Spammers and rule-breakers do the same thing to automated filters. They spread the letters out, or swap a digit for a similar-looking letter. The human reader still gets the message. A naive keyword check sees nothing at all.
Why it exists
Simple content filters check whether a banned word appears in the text, as an exact string. That check is easy to defeat once someone knows it exists. Insert a space, swap "o" for "0," and the exact string is gone.
This is not usually a first-time accident. It is a deliberate, adversarial move by someone who already knows a filter is watching. Evasion-resistant filtering exists because of this. A naive filter and a determined evader stay in an active back-and-forth. It is not a one-time problem to solve.
How it works
"f r e e m0ney"
|
v
remove spacing tricks, map lookalike characters back
|
v
"freemoney"
|
v
now the banned-word check actually matchesWhere you have already seen it
- Spam email with oddly spaced or symbol-substituted words, specifically built to dodge spam filters.
- Social media comments using creative spelling to discuss banned or sensitive topics without tripping moderation.
- Online game chat, where players invent new spellings for blocked words within days of a filter update.
Remember this
- Filter evasion is a deliberate, adversarial rewrite of a word, not an accidental typo.
- Normalising spacing tricks and common character substitutions before checking against a banned-word list catches much of it.
- This is a moving target. Evasion techniques adapt as soon as a filter updates. No fixed rule set stays effective forever.
What to learn next
- Cleaning social media text — the related, non-adversarial cleanup this lesson's technique builds on.
- Text classification — the broader moderation task this kind of filter usually supports.
- Finding and masking personal data in text — another rule-based detection task, defending against a very different kind of risk.
Developer — Code and libraries.
Setup
Nothing to install. Pure Python standard library — re.
Normalising common evasion tricks before filtering
import re
LEET_MAP = str.maketrans({"0": "o", "1": "i", "3": "e", "4": "a", "5": "s", "@": "a", "$": "s"})
def normalise_evasion(text):
text = text.lower()
text = text.translate(LEET_MAP)
text = re.sub(r"[\s._\-*]+", "", text) # collapse spacing tricks
return text
BANNED = ["freemoney", "winnow", "clickhere"]
samples = ["f r e e m0ney", "FREE.MONEY!!!", "fr33 m0n3y", "wi n_n0w today"]
for s in samples:
cleaned = normalise_evasion(s)
hit = any(b in cleaned for b in BANNED)
print(f"{s!r:20} -> {cleaned!r:16} flagged={hit}")'f r e e m0ney' -> 'freemoney' flagged=True 'FREE.MONEY!!!' -> 'freemoney!!!' flagged=True 'fr33 m0n3y' -> 'freemoney' flagged=True 'wi n_n0w today' -> 'winnowtoday' flagged=True
Line by line
LEET_MAP reverses common lookalike-character substitutions — "0" back to "o," "3" back to "e," "@" back to "a" — the specific character swaps sometimes called "leetspeak."
Removing whitespace, periods, underscores and dashes catches spacing-based evasion. "f r e e m0ney" has spaces inserted between every letter — stripping them collapses it right back to "freemoney," which the banned-word check then matches directly.
The banned-word check is a plain substring test, run only after normalisation. any(b in cleaned for b in BANNED) — deliberately simple, since all the actual evasion-defeating work already happened in normalise_evasion.
"wi n_n0w today" flagged as containing "winnow," which is technically correct but worth noticing. The cleaned string "winnowtoday" does contain "winnow" as a substring — the filter is working exactly as written, even though a human moderator would likely judge this specific message as harmless. This is a real trade-off this exact demo makes visible: aggressive normalisation catches more evasion, and also raises the false-positive rate.
Common mistakes
Stripping spaces before checking for spacing-based evasion at all. The order in the function matters — case-folding and character-mapping happen first, spacing removal happens last, so that mapped characters are correctly lowercase before the final substring check runs.
Building a leetspeak map that is too permissive. Every character substitution added to LEET_MAP also increases false-positive risk on legitimate text — "s3rver" mapping to "server" might be desired, but an overly broad map starts corrupting ordinary technical vocabulary that never meant to evade anything.
Assuming this fixed rule set stays effective indefinitely. The moment evaders learn a specific normalisation rule exists, they route around it — using a Cyrillic lookalike character instead of a digit, for instance, which nothing in LEET_MAP accounts for.
Try it yourself
Add a Unicode homoglyph to LEET_MAP — for example, mapping Cyrillic "а" (which looks identical to Latin "a" but is a different character) back to Latin "a." Test it against a sample that uses the Cyrillic character deliberately, and see whether the check still catches it. This is a real, actively-used evasion technique, well beyond simple digit substitution.
What to learn next
- Cleaning social media text — the non-adversarial version of the normalisation used here.
- Regular expressions for text work — a fuller treatment of the pattern-matching tools used throughout this lesson.
- Finding and masking personal data in text — a different rule-based defensive filter, against a different kind of risk.
Researcher — Mathematics and papers.
Filter evasion as an adversarial game, not a static classification problem
Unlike most text-processing tasks covered in this section, filter evasion involves an active adversary who observes and adapts to the filter's known behaviour. This changes the right framing: a fixed rule set, however comprehensive at deployment time, has a shelf life determined by how quickly evaders discover and route around its specific weaknesses — a dynamic closer to spam-filtering's long adversarial history than to a stationary classification task with a fixed, unchanging data distribution.
Categorising evasion techniques
Character-level substitution. Digit-for-letter (leetspeak), Unicode homoglyphs (visually identical characters from different scripts — Cyrillic "а" for Latin "a"), and diacritic insertion, all targeting exact-string or simple-regex matching directly.
Whitespace and delimiter insertion. Spaces, periods, zero-width characters (including the genuinely invisible zero-width space, U+200B) inserted between letters of a banned word, defeating substring matching without changing how a human reader perceives the word.
Semantic evasion. Replacing a banned term with an approved synonym, a coded euphemism, or a community-specific slang term carrying the same meaning — this class defeats any normalisation-based approach, since there is no shared surface form to normalise toward at all. It requires meaning-level classification, not string matching, and is the primary reason modern moderation systems combine rule-based prefiltering with a trained classifier rather than relying on either alone.
Formal treatment: normalisation as an equivalence class
The developer demo's normalise_evasion defines an equivalence relation: two strings are treated as equivalent if they normalise to the same canonical form.
s ~ t iff normalise(s) = normalise(t)A well-designed normalisation function should be sound (only merges strings a human would judge as intentional variants of the same underlying word) and reasonably complete (catches the evasion patterns actually observed in practice). These two goals trade off directly against each other: a broader normalisation function catches more evasion at the cost of more false-positive merges on legitimate text, exactly the "winnow" versus "winnowtoday" tension the developer demo surfaces directly.
Key references
- Yih, W., Goodman, J. & Hulten, G. (2006). Learning at Low False Positive Rates. CEAS. — early formal treatment of the precision-recall trade-off in adversarial spam filtering.
- Fabbri, F. et al. (2013). A Survey on Adversarial Text Classification. — evasion techniques and defences across the broader adversarial text classification literature.
- Le, T. et al. (2022). Perturbations in the Wild: Leveraging Human-Written Text Perturbations for Realistic Adversarial Attack and Defense. ACL Findings. — catalogues real-world evasion patterns observed against deployed content moderation systems.
Current state and open problems
Character-level normalisation, as demonstrated above, remains a cheap and genuinely useful first-pass filter against the character-substitution and spacing categories of evasion, and is standard practice at essentially every platform with automated moderation.
Semantic evasion is the unresolved frontier: coded language and community-specific euphemisms shift faster than any fixed classifier retraining cycle can track, and detecting them reliably requires either continuously refreshed training data reflecting current evasion patterns, or a much stronger contextual understanding of intent than keyword-adjacent classifiers currently provide. This remains an active, unsolved arms race rather than a task with a stable, final solution.
What to learn next
- Cleaning social media text — the shared normalisation techniques this lesson repurposes for adversarial defence.
- Text classification — the meaning-level classification approach needed once semantic evasion defeats normalisation.
- Finding and masking personal data in text — another rule-based text-defence task, though against a non-adversarial risk.