Cleaning social media text
Social media text is full of hashtags, mentions, stretched-out letters and links, and cleaning it means removing the noise without losing the meaning.
- 8 min read
- 3 reading levels
- Published
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Cleaning social media text strips out the noise — links, mentions, stretched-out letters — while keeping the actual meaning intact.
Think about reading a friend's excited voice note. "Sooooo good!!!" carries real enthusiasm. If you were writing down only the facts for someone else, you would note "so good" and move on. You filtered the noise, kept the meaning.
Cleaning social media text does the same job automatically: strip the decoration, keep what the post is actually about.
Why it exists
Posts, comments and reviews are written fast, casually, for other people to read — not for a language model. They carry links that add nothing to the text's meaning. They carry mentions of other users, and hashtags doing double duty as content and metadata. Letters get stretched out for emphasis.
Feed that straight into most NLP tools and the noise drowns the signal. "sooooo" and "so" look like unrelated words, unless something was taught that they mean the same thing.
How it works
"omg this movieeeee was sooooo good!!! @friend #bollywood http://link.com"
| | | | |
keep collapse collapse keep as strip
word repeats repeats a tag (adds no meaning)Each kind of noise gets its own small rule. Strip links entirely. Pull hashtags out as a separate signal. Collapse letters repeated more than twice. Remove mentions.
Where you have already seen it
- Sentiment analysis tools built on top of cleaned tweets and reviews, working on the meaning rather than the decoration.
- Trending topic detection, extracting hashtags as their own signal, separate from the sentence they appeared in.
- Search and indexing on social platforms, where a raw link in a post would otherwise pollute keyword matching.
Remember this
- Social media noise — links, mentions, stretched letters, hashtags — has predictable patterns, and each pattern gets its own cleaning rule.
- Collapsing repeated letters helps a model see "so" and "sooooo" as related. The cost: it sometimes over-collapses legitimate double letters.
- Hashtags carry real content and are usually extracted separately, not deleted outright.
What to learn next
- Fixing OCR errors — a different, non-social source of messy text.
- Regular expressions for text work — the pattern-matching tool every rule in this lesson is built from.
- When users write 'f r e e m0ney' to dodge your filter — messy text, made that way on purpose.
Developer — Code and libraries.
Setup
Nothing to install. Pure Python standard library — re.
Cleaning a tweet-shaped piece of text
import re
def clean_tweet(text):
text = re.sub(r"http\S+|www\.\S+", "", text) # links
text = re.sub(r"@\w+", "", text) # mentions
hashtags = re.findall(r"#(\w+)", text)
text = re.sub(r"#\w+", "", text) # strip hashtags, keep words separately
text = re.sub(r"(.)\1{2,}", r"\1\1", text) # sooooo -> soo (cap repeats at 2)
text = re.sub(r"\s+", " ", text).strip()
return text, hashtags
samples = [
"omg this movieeeee was sooooo good!!! @friend check it out #bollywood #mustwatch http://example.com",
"just landed in Bengaluru soooo tired but happpppy",
]
for s in samples:
cleaned, tags = clean_tweet(s)
print("IN: ", s)
print("OUT:", cleaned, "| tags:", tags)
print()IN: omg this movieeeee was sooooo good!!! @friend check it out #bollywood #mustwatch http://example.com OUT: omg this moviee was soo good!! check it out | tags: ['bollywood', 'mustwatch'] IN: just landed in Bengaluru soooo tired but happpppy OUT: just landed in Bengaluru soo tired but happy | tags: []
Line by line
Links and mentions get removed entirely — http\S+ matches a URL until the next whitespace, @\w+ matches a mention. Neither carries meaning useful to most downstream tasks, so both are dropped rather than transformed.
Hashtags are pulled out with findall before being stripped from the text. #bollywood and #mustwatch become their own list, ['bollywood', 'mustwatch'] — genuinely useful topic signal, kept separately rather than lost.
(.)\1{2,} finds any character repeated three or more times in a row, and collapses it to two. "sooooo" (five o's) becomes "soo" (two o's) — enough to normalise stretched-out emphasis, without collapsing all the way to a single letter and losing the fact that emphasis happened at all.
Look closely at "movieeeee" becoming "moviee." The word "movie" legitimately has one "e," and the stretched ending adds four more — the regex, correctly following its rule, still leaves two e's behind instead of one, since it caps repeats at two rather than one. This is expected behaviour, not a bug, though it means the cleaned word is not always a real dictionary word.
Common mistakes
Collapsing repeats down to a single character instead of two. "committee," "bookkeeper," and many other real words contain legitimate double letters. Capping at two rather than one avoids mangling ordinary words that happen to have doubled letters, though it does not fix every case, as "moviee" shows above.
Deleting hashtags instead of extracting them first. A hashtag is content wearing a metadata costume — "#mustwatch" is a real opinion signal for sentiment analysis. Throwing it away entirely discards information a careful pipeline would keep.
Running URL-stripping after whitespace collapsing, in the wrong order. Order matters here: strip links and mentions first, then collapse the leftover multiple spaces they left behind — reversing the order can leave stray double spaces around where a link used to be.
Try it yourself
Add a sample with an emoji in it, and check whether clean_tweet leaves it untouched (it does — nothing in this function targets emoji specifically). Decide whether your downstream task wants emoji kept, stripped, or converted to words like ":smiling_face:", and add a rule for whichever choice fits your case.
What to learn next
- When users write 'f r e e m0ney' to dodge your filter — text deliberately obfuscated, rather than written casually by accident.
- Regular expressions for text work — a deeper look at the pattern language used throughout this lesson.
- Stopwords: when removing them helps and when it breaks things — a further cleanup step commonly applied after this one.
Researcher — Mathematics and papers.
Framing this as a rule-based normalisation pipeline
Social media cleaning is a composition of independent, order-sensitive regular-expression substitutions, each targeting one well-defined noise pattern:
clean(x) = f_n( f_{n-1}( ... f_1(x) ... ) )- Each
f_iis a regex substitution or extraction targeting one noise category — URLs, mentions, character elongation, and so on. - Order matters when patterns can interact: elongation-collapsing before whitespace-collapsing, link-stripping before elongation-collapsing (a URL's repeated characters, such as
www, should not be treated as emphasis).
This pipeline shape is standard across production social-text preprocessing (Baziotis et al., 2017, the ekphrasis library documents a comparable rule ordering) and remains dominant over learned approaches for this specific normalisation layer, since the noise patterns are well-understood, stable, and cheap to express as rules.
Character elongation as a linguistic signal, not only noise
Elongation ("soooo," "nooooo") is not random — sociolinguistic research on computer-mediated communication treats it as a genuine paralinguistic device signalling emphasis or emotional intensity, roughly analogous to prosodic stress in speech (Brody & Diakopoulos, 2011, on lengthening as a sentiment-bearing signal in social text specifically). Naively collapsing all elongation to the same fixed form, as the developer demo does, discards this intensity signal. Sentiment-aware preprocessing sometimes preserves elongation length as a separate numeric feature — "how many extra characters" — rather than normalising it away entirely.
Hashtag segmentation
findall(r"#(\w+)") extracts a hashtag as one unsegmented token — #mustwatch becomes the string "mustwatch," not the words "must" and "watch." Camel-case hashtags (#MustWatchNow) can be split on capitalisation boundaries directly; all-lowercase concatenated hashtags (#mustwatchnow) require a word-segmentation algorithm, commonly a Viterbi search over a word-frequency lexicon (Wang et al., 2011, on Twitter-specific hashtag segmentation), similar in spirit to the maximum-likelihood segmentation approach covered for Unigram tokenization in tokenization.
Key references
- Baziotis, C., Pelekis, N. & Doulkeridis, C. (2017). DataStories at SemEval-2017 Task 4: Deep LSTM with Attention for Message-level and Topic-based Sentiment Analysis. SemEval. — describes the
ekphrasispreprocessing pipeline widely used as a reference implementation for this task. - Brody, S. & Diakopoulos, N. (2011). Cooooooooooooooollllllllllllll!!!!!!!!!!!!!! Using Word Lengthening to Detect Sentiment in Microblogs. EMNLP.
- Wang, K. et al. (2011). A Fast Algorithm for Segmenting Hashtags into Words. — Viterbi-based hashtag segmentation using a lexicon-and-frequency scoring approach.
Current state and open problems
Emoji present the same "signal, not noise" problem elongation does, and are entirely untouched by the developer demo above. Production systems commonly map emoji to a text description (via a fixed lookup table) rather than stripping them, since emoji frequently carry sentiment or intent that the surrounding text does not restate — a laughing-face emoji can flip the polarity of an otherwise neutral sentence.
Platform-specific conventions keep shifting — hashtag usage, mention syntax and elongation patterns differ meaningfully between platforms and change over time within one platform — meaning a cleaning pipeline tuned for one dataset degrades gracefully but noticeably when applied to a different platform's text without re-tuning.
What to learn next
- When users write 'f r e e m0ney' to dodge your filter — deliberate obfuscation, contrasted against the casual noise covered here.
- Regular expressions for text work — the full pattern-matching toolkit this pipeline is built from.
- Tokenization — the segmentation problem hashtag-splitting is a smaller version of.