Messy Real-World Text

Augmenting text data

Text data augmentation creates new, slightly varied training examples from existing ones, so a model sees more variety without needing more labelled data.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Text data augmentation creates new training examples by slightly changing existing ones. A model gets more variety, without collecting more real data.

Think about practising a speech. You do not say the exact same sentence every time you rehearse. You swap a word here, reorder a phrase there — roughly the same thing, said slightly differently. Each version is still recognisably the same speech. Rehearsing several versions makes you more adaptable on the actual day.

Data augmentation gives a model that same varied practice: the same underlying example, expressed several slightly different ways.

Why it exists

Labelled training data is expensive to collect — someone has to write or label each example by hand. Small labelled datasets lead models to memorise the exact wording they were shown, instead of learning the underlying pattern.

Augmentation manufactures extra, varied examples from the data already collected. A model trained on an original sentence, plus several slightly reworded versions, learns the pattern behind the wording. It does not memorise the exact words used.

How it works

  original:  "the food was good and the service was fast"
                          |
          ---------------------------------
          |                               |
          v                               v
   "the food was great        "the service was good
    and the service            and the food was fast"
    was fast"                   (word order shuffled)
   (synonym swapped)

Where you have already seen it

  • Image tools rotate and crop photos, teaching a vision model the same object from different angles.
  • Chatbot training data is often expanded with paraphrases, so the bot recognises a request phrased several ways.
  • Spam filters trained on augmented text, catching a spam message even after a spammer changes a few words.

Remember this

  • Augmentation manufactures extra training examples from existing labelled data, without collecting anything new.
  • It works because a model shown only one wording per idea tends to memorise wording, not meaning.
  • Every augmentation technique risks changing the meaning by accident. A synonym swap must actually preserve meaning, or the new example is wrong.

What to learn next

Developer — Code and libraries.

Setup

Nothing to install. Pure Python standard library — random.

Two classic augmentation techniques

text_augment_demo.py
import random
random.seed(7)

# A tiny hand-built synonym table. Real pipelines use WordNet or an embedding model.
SYNONYMS = {
    "good": ["great", "nice", "solid"],
    "bad": ["poor", "weak", "awful"],
    "fast": ["quick", "speedy", "rapid"],
    "happy": ["pleased", "glad", "content"],
    "big": ["large", "huge", "sizeable"],
}

def synonym_replace(text, n=1):
    words = text.split()
    candidates = [i for i, w in enumerate(words) if w.lower() in SYNONYMS]
    random.shuffle(candidates)
    for i in candidates[:n]:
        words[i] = random.choice(SYNONYMS[words[i].lower()])
    return " ".join(words)

def random_swap(text, n=1):
    words = text.split()
    for _ in range(n):
        if len(words) < 2:
            break
        i, j = random.sample(range(len(words)), 2)
        words[i], words[j] = words[j], words[i]
    return " ".join(words)

original = "the food was good and the service was fast"
for _ in range(3):
    print("synonym:", synonym_replace(original))
for _ in range(3):
    print("swap:   ", random_swap(original))
Output
synonym: the food was great and the service was fast
synonym: the food was solid and the service was fast
synonym: the food was good and the service was quick
swap:    the fast was good and the service was food
swap:    the food was good and the service was fast
swap:    the food was fast and the service was good

This output is exact and repeatable — random.seed(7) fixes every random choice. A different seed, or no seed at all, would produce different specific swaps each run, though the technique behaves the same way.

Line by line

synonym_replace only ever touches words present in the SYNONYMS table. This keeps the augmentation safe and predictable — it cannot accidentally introduce a nonsense word, only a genuine synonym of one it already recognises.

random_swap can produce genuinely broken sentences. "the fast was good and the service was food" is grammatically damaged, and reads as nonsense. This is a real, known weakness of random swapping — it helps a model tolerate word-order noise, at the cost of sometimes training on a sentence that no longer reliably means what the label says it means.

Every call uses Python's shared random state, seeded once at the top of the script. That single seed is what makes every run of this exact script produce the exact output block above.

Common mistakes

Swapping or replacing a word that changes the label. "the food was good" augmented into "the food was bad" via a careless synonym table entry would flip a positive review into a negative one — the augmented example would then actively teach the model something false.

Over-augmenting a small dataset with the same few techniques repeatedly. Ten augmented variants of one original sentence are still fundamentally one data point in disguise — augmentation adds variety around existing examples, it does not manufacture genuinely new information.

Using a tiny hand-built synonym table like this one in production. It covers five words. A real pipeline uses WordNet, a thesaurus API, or an embedding-based nearest-neighbour lookup to cover realistic vocabulary size — this demo's table is deliberately minimal, to keep every step inspectable.

Try it yourself

Add a random_delete(text, p) function that drops each word with probability p, another classic augmentation technique (Wei & Zou, 2019, call this part of "Easy Data Augmentation"). Run it on the same sentence several times with a fixed seed, and check how much information a sentence can lose before it stops being a useful training example at all.

What to learn next

Researcher — Mathematics and papers.

A taxonomy of text augmentation techniques

Word-level, rule-based. Synonym replacement, random swap, random deletion, random insertion — collectively "Easy Data Augmentation" (Wei & Zou, 2019). Cheap, fast, no model required. Reported gains are largest on small training sets (hundreds of examples) and shrink toward negligible as dataset size grows into the thousands, since a large enough real dataset already captures the variety these techniques approximate.

Back-translation (Sennrich, Haddow & Birch, 2016; Edunov et al., 2018, for its large-scale application). Translate an example to another language and back, producing a fluent paraphrase that differs in surface wording while preserving meaning far more reliably than word-level rules, at the cost of needing a translation model — see translating a low-resource language.

Contextual embedding substitution (Kobayashi, 2018). Replace a word with one sampled from a masked language model's predicted distribution at that position, conditioned on surrounding context — more fluent and context-appropriate than a static synonym table, since the replacement is chosen given the actual sentence rather than from a fixed word-to-word mapping.

Mixup and interpolation-based methods (Guo, Mao & Zhang, 2019, adapting Zhang et al.'s 2017 image-domain Mixup to text). Interpolate between two examples' hidden representations and their labels, rather than manipulating text directly — a fundamentally different mechanism from the surface-level methods above, operating in embedding space instead of on the text itself.

Generative augmentation. Prompting a large language model to produce paraphrases or entirely new labelled examples matching a given pattern — increasingly the dominant approach where API or local LLM access is available, since it can produce more fluent and varied output than rule-based methods without needing a dedicated translation or masked-LM pipeline.

Why augmentation helps: a regularisation view

Augmentation is a form of regularisation via data-space perturbation. Given a base training distribution and an augmentation function producing perturbed variants near each real example, training on the augmented set encourages the model's decision boundary to stay stable within a local neighbourhood of each training point, rather than fitting tightly to the exact wording observed.

text
loss_augmented = E over (x, y) in D, x' ~ augment(x) of  loss( f(x'), y )
  • augment(x) is the augmentation function's output distribution given input x — the specific transform (synonym swap, back-translation, embedding substitution) determines what "nearby" means for that method.
  • This is the text-domain analogue of image augmentation (rotation, crop, colour jitter), sharing the same underlying justification: better generalisation from a fixed labelled dataset, without collecting more raw data.

Key references

  • Wei, J. & Zou, K. (2019). EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks. arXiv:1901.11196
  • Sennrich, R., Haddow, B. & Birch, A. (2016). Improving Neural Machine Translation Models with Monolingual Data. arXiv:1511.06709 — back-translation.
  • Kobayashi, S. (2018). Contextual Augmentation: Data Augmentation by Words with Paradigmatic Relations. arXiv:1805.06201
  • Guo, H., Mao, Y. & Zhang, R. (2019). Augmenting Data with Mixup for Sentence Classification: An Empirical Study. arXiv:1905.08941
  • Feng, S. et al. (2021). A Survey of Data Augmentation Approaches for NLP. arXiv:2105.03075 — a comprehensive overview spanning all the technique families above.

Current state and open problems

Generative augmentation via prompted LLMs has, for most practical purposes, displaced hand-coded rule-based methods where API or local model access is available, since it produces more fluent, more varied, and more reliably meaning-preserving output than synonym tables or random word operations.

The unresolved question is quality control: an LLM-generated "augmented" example can still drift from the original label — the same risk word-level swapping carries, now harder to spot by eye because the output reads fluently even when subtly wrong. Automated filtering of generated augmentations, checking that the augmented example still plausibly carries the original label, is an active area rather than a solved step in the pipeline — see labelling data with an LLM and a human reviewer for the closely related problem of trusting a model's own output before it enters a training set.

What to learn next