Messy Real-World Text

Cleaning social media text

Social media text is full of hashtags, mentions, stretched-out letters and links, and cleaning it means removing the noise without losing the meaning.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Cleaning social media text strips out the noise — links, mentions, stretched-out letters — while keeping the actual meaning intact.

Think about reading a friend's excited voice note. "Sooooo good!!!" carries real enthusiasm. If you were writing down only the facts for someone else, you would note "so good" and move on. You filtered the noise, kept the meaning.

Cleaning social media text does the same job automatically: strip the decoration, keep what the post is actually about.

Why it exists

Posts, comments and reviews are written fast, casually, for other people to read — not for a language model. They carry links that add nothing to the text's meaning. They carry mentions of other users, and hashtags doing double duty as content and metadata. Letters get stretched out for emphasis.

Feed that straight into most NLP tools and the noise drowns the signal. "sooooo" and "so" look like unrelated words, unless something was taught that they mean the same thing.

How it works

  "omg this movieeeee was sooooo good!!! @friend #bollywood http://link.com"
       |            |          |                |         |
     keep        collapse   collapse          keep as    strip
     word        repeats    repeats            a tag    (adds no meaning)

Each kind of noise gets its own small rule. Strip links entirely. Pull hashtags out as a separate signal. Collapse letters repeated more than twice. Remove mentions.

Where you have already seen it

  • Sentiment analysis tools built on top of cleaned tweets and reviews, working on the meaning rather than the decoration.
  • Trending topic detection, extracting hashtags as their own signal, separate from the sentence they appeared in.
  • Search and indexing on social platforms, where a raw link in a post would otherwise pollute keyword matching.

Remember this

  • Social media noise — links, mentions, stretched letters, hashtags — has predictable patterns, and each pattern gets its own cleaning rule.
  • Collapsing repeated letters helps a model see "so" and "sooooo" as related. The cost: it sometimes over-collapses legitimate double letters.
  • Hashtags carry real content and are usually extracted separately, not deleted outright.

What to learn next

Developer — Code and libraries.

Setup

Nothing to install. Pure Python standard library — re.

Cleaning a tweet-shaped piece of text

clean_social_text.py
import re

def clean_tweet(text):
    text = re.sub(r"http\S+|www\.\S+", "", text)              # links
    text = re.sub(r"@\w+", "", text)                          # mentions
    hashtags = re.findall(r"#(\w+)", text)
    text = re.sub(r"#\w+", "", text)                          # strip hashtags, keep words separately
    text = re.sub(r"(.)\1{2,}", r"\1\1", text)                 # sooooo -> soo (cap repeats at 2)
    text = re.sub(r"\s+", " ", text).strip()
    return text, hashtags

samples = [
    "omg this movieeeee was sooooo good!!! @friend check it out #bollywood #mustwatch http://example.com",
    "just landed in Bengaluru soooo tired but happpppy",
]
for s in samples:
    cleaned, tags = clean_tweet(s)
    print("IN: ", s)
    print("OUT:", cleaned, "| tags:", tags)
    print()
Output
IN:  omg this movieeeee was sooooo good!!! @friend check it out #bollywood #mustwatch http://example.com
OUT: omg this moviee was soo good!! check it out | tags: ['bollywood', 'mustwatch']

IN:  just landed in Bengaluru soooo tired but happpppy
OUT: just landed in Bengaluru soo tired but happy | tags: []

Line by line

Links and mentions get removed entirely — http\S+ matches a URL until the next whitespace, @\w+ matches a mention. Neither carries meaning useful to most downstream tasks, so both are dropped rather than transformed.

Hashtags are pulled out with findall before being stripped from the text. #bollywood and #mustwatch become their own list, ['bollywood', 'mustwatch'] — genuinely useful topic signal, kept separately rather than lost.

(.)\1{2,} finds any character repeated three or more times in a row, and collapses it to two. "sooooo" (five o's) becomes "soo" (two o's) — enough to normalise stretched-out emphasis, without collapsing all the way to a single letter and losing the fact that emphasis happened at all.

Look closely at "movieeeee" becoming "moviee." The word "movie" legitimately has one "e," and the stretched ending adds four more — the regex, correctly following its rule, still leaves two e's behind instead of one, since it caps repeats at two rather than one. This is expected behaviour, not a bug, though it means the cleaned word is not always a real dictionary word.

Common mistakes

Collapsing repeats down to a single character instead of two. "committee," "bookkeeper," and many other real words contain legitimate double letters. Capping at two rather than one avoids mangling ordinary words that happen to have doubled letters, though it does not fix every case, as "moviee" shows above.

Deleting hashtags instead of extracting them first. A hashtag is content wearing a metadata costume — "#mustwatch" is a real opinion signal for sentiment analysis. Throwing it away entirely discards information a careful pipeline would keep.

Running URL-stripping after whitespace collapsing, in the wrong order. Order matters here: strip links and mentions first, then collapse the leftover multiple spaces they left behind — reversing the order can leave stray double spaces around where a link used to be.

Try it yourself

Add a sample with an emoji in it, and check whether clean_tweet leaves it untouched (it does — nothing in this function targets emoji specifically). Decide whether your downstream task wants emoji kept, stripped, or converted to words like ":smiling_face:", and add a rule for whichever choice fits your case.

What to learn next

Researcher — Mathematics and papers.

Framing this as a rule-based normalisation pipeline

Social media cleaning is a composition of independent, order-sensitive regular-expression substitutions, each targeting one well-defined noise pattern:

text
clean(x) = f_n( f_{n-1}( ... f_1(x) ... ) )
  • Each f_i is a regex substitution or extraction targeting one noise category — URLs, mentions, character elongation, and so on.
  • Order matters when patterns can interact: elongation-collapsing before whitespace-collapsing, link-stripping before elongation-collapsing (a URL's repeated characters, such as www, should not be treated as emphasis).

This pipeline shape is standard across production social-text preprocessing (Baziotis et al., 2017, the ekphrasis library documents a comparable rule ordering) and remains dominant over learned approaches for this specific normalisation layer, since the noise patterns are well-understood, stable, and cheap to express as rules.

Character elongation as a linguistic signal, not only noise

Elongation ("soooo," "nooooo") is not random — sociolinguistic research on computer-mediated communication treats it as a genuine paralinguistic device signalling emphasis or emotional intensity, roughly analogous to prosodic stress in speech (Brody & Diakopoulos, 2011, on lengthening as a sentiment-bearing signal in social text specifically). Naively collapsing all elongation to the same fixed form, as the developer demo does, discards this intensity signal. Sentiment-aware preprocessing sometimes preserves elongation length as a separate numeric feature — "how many extra characters" — rather than normalising it away entirely.

Hashtag segmentation

findall(r"#(\w+)") extracts a hashtag as one unsegmented token — #mustwatch becomes the string "mustwatch," not the words "must" and "watch." Camel-case hashtags (#MustWatchNow) can be split on capitalisation boundaries directly; all-lowercase concatenated hashtags (#mustwatchnow) require a word-segmentation algorithm, commonly a Viterbi search over a word-frequency lexicon (Wang et al., 2011, on Twitter-specific hashtag segmentation), similar in spirit to the maximum-likelihood segmentation approach covered for Unigram tokenization in tokenization.

Key references

  • Baziotis, C., Pelekis, N. & Doulkeridis, C. (2017). DataStories at SemEval-2017 Task 4: Deep LSTM with Attention for Message-level and Topic-based Sentiment Analysis. SemEval. — describes the ekphrasis preprocessing pipeline widely used as a reference implementation for this task.
  • Brody, S. & Diakopoulos, N. (2011). Cooooooooooooooollllllllllllll!!!!!!!!!!!!!! Using Word Lengthening to Detect Sentiment in Microblogs. EMNLP.
  • Wang, K. et al. (2011). A Fast Algorithm for Segmenting Hashtags into Words. — Viterbi-based hashtag segmentation using a lexicon-and-frequency scoring approach.

Current state and open problems

Emoji present the same "signal, not noise" problem elongation does, and are entirely untouched by the developer demo above. Production systems commonly map emoji to a text description (via a fixed lookup table) rather than stripping them, since emoji frequently carry sentiment or intent that the surrounding text does not restate — a laughing-face emoji can flip the polarity of an otherwise neutral sentence.

Platform-specific conventions keep shifting — hashtag usage, mention syntax and elongation patterns differ meaningfully between platforms and change over time within one platform — meaning a cleaning pipeline tuned for one dataset degrades gracefully but noticeably when applied to a different platform's text without re-tuning.

What to learn next