Sentiment, Opinion and Text Mining
Toxicity and abuse detection
Toxicity detection scores whether text is abusive, threatening or hateful, usually across several separate categories rather than one single flag.
- 8 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Toxicity detection scores whether text is abusive, threatening or hateful, so it can be moderated before it reaches other people.
Picture a bouncer at a venue door, checking every person before they walk in. Most people pass through without a second look. A few get stopped, for a specific, checkable reason.
Toxicity detection works the same way for text. Every comment gets scanned. Most pass through untouched. A few get flagged, scored against specific categories like insults, threats or hate speech.
Why it exists
Any platform where strangers can post text needs a way to catch abuse before it reaches other users. No team of human moderators alone could review it at that scale.
A single "toxic or not" flag is often too blunt. A platform needs to know whether a comment is only rude, actually threatening, or targets someone's identity. Each of those needs a different response.
Toxicity models exist to do this scoring automatically. They run fast enough for every single post, with specific categories a moderation team can act on.
How it works
"I disagree with your point, but you make a fair argument."
-> toxic: 0.001 (essentially clean)
"You are an absolute idiot and everyone here hates you."
-> toxic: 0.992, insult: 0.949, obscene: 0.695A toxicity model scores text across several categories at once, not only one. A comment can be highly insulting without being obscene, or threatening without either.
Where you have already seen it
- Comment sections that hide a reply behind "this comment may be offensive." A toxicity score decided that, before any human read it.
- YouTube and Twitch chat filters during a livestream. Scoring messages fast enough to filter them in real time.
- Customer support tools that flag abusive messages for a supervisor. Routing based on toxicity, not only general sentiment.
- Community moderation dashboards that prioritise the worst content first. Toxicity score decides the review queue order.
Remember this
- Toxicity detection scores text across specific categories, not only one overall flag.
- It exists to catch abuse at a scale humans cannot review alone.
- A comment can score high on one category and low on another; they are not the same thing.
What to learn next
- Fine-grained sentiment — a related but different kind of scoring, not focused on harm.
- Bias in datasets — a real risk in exactly this kind of model, covered next.
- Emotion classification — scoring feeling rather than harm.
Developer — Code and libraries.
Below, a model trained on labelled toxic comments scores two messages across several categories at once.
Setup
pip install transformers torchThe first run downloads unitary/toxic-bert, about 440 MB.
Scoring two messages
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="unitary/toxic-bert",
top_k=None,
)
texts = [
"I disagree with your point, but I think you make a fair argument overall.",
"You are an absolute idiot and everyone here hates you.",
]
for text in texts:
scores = classifier(text)[0]
scores.sort(key=lambda x: x["score"], reverse=True)
print(text)
for s in scores[:3]:
print(f" {s['label']:15s} {s['score']:.3f}")
print()I disagree with your point, but I think you make a fair argument overall. toxic 0.001 obscene 0.000 insult 0.000 You are an absolute idiot and everyone here hates you. toxic 0.992 insult 0.949 obscene 0.695
Line by line
top_k=None returns a score for every category the model knows, not only the top one. This model reports several labels, including toxic, insult and obscene, each scored independently.
The disagreement message scored close to zero across the board. Disagreement itself is not toxicity. The model correctly separated "pushing back on an argument" from "attacking a person."
The insult message scored high on toxic and insult, but noticeably lower on obscene. No profanity appears in the sentence at all. "Obscene" at 0.695 is likely picking up on the harsh, degrading tone rather than any specific word.
Common mistakes
Using a single toxicity threshold for every category. A 0.5 cutoff that works well for "obscene" may be too strict or too loose for "threat." Categories differ in base rate, and in how confidently the model scores them.
Auto-deleting content on a toxicity score alone, with no appeal path. These models make mistakes, including false positives on strong but non-abusive language. Any high-stakes moderation action needs a review path, not a silent automatic removal.
Ignoring known bias in toxicity models. Several published toxicity classifiers score text mentioning certain identity terms as more toxic on average, even in neutral sentences. This happens purely from how the training data got collected. Always test a toxicity model against neutral sentences naming different identities before trusting it in production.
Assuming toxicity transfers across languages and platforms. A model trained on English comment data will not reliably score Hindi, Hinglish, or heavily platform-specific slang the same way. Test on your actual traffic, not assumptions.
Try it yourself
Try a sentence that is blunt but not abusive, like "This idea is genuinely bad and should not go forward." Compare its scores against the plainly insulting example above, and think about where you would draw a moderation line.
What to learn next
- Bias in datasets — the identity-term bias problem named directly above.
- Choosing a threshold from costs — deciding where a toxicity cutoff should actually sit.
- Fairness metrics — measuring whether a model treats different groups consistently.
Researcher — Mathematics and papers.
Multi-label toxicity as the standard framing
Toxicity detection is typically framed as multi-label binary classification, across a fixed taxonomy: toxic, severe_toxic, obscene, threat, insult, identity_hate. The widely used Jigsaw Toxic Comment dataset (cjadams et al., 2017) trains models like the one in the developer block. Each label is predicted independently. That is why a message can score high on insult while low on obscene, exactly as in the developer block.
The identity-term bias problem
Dixon et al. (2018) document a specific, measurable failure. Models trained on Jigsaw-style data score neutral sentences mentioning certain identity terms as more toxic. Religion, sexual orientation and disability are typical examples, compared against matched sentences without those terms. This arises from a base-rate confound in training data. Comments mentioning these terms co-occur with toxic comments more often in the raw corpus. The reason is unrelated to the term itself. Those terms co-occurred, in the training data, with discussions that attracted more abuse.
Borkan et al. (2019) formalise this as unintended bias. They propose subgroup-specific evaluation metrics, comparing model performance on sentences mentioning a given identity term against overall performance. This surfaces the failure during evaluation, rather than after deployment.
Adversarial evasion
Toxic content evolves specifically to evade detection. Leetspeak substitutions, added spacing, and homoglyphs all aim to fool a classifier. Platform-specific coded language does the same, while staying readable to a human. This is covered directly in when users write 'f r e e m0ney' to dodge your filter. Toxicity models trained on clean historical data degrade measurably against these evasion techniques, requiring ongoing retraining as evasion patterns shift.
Context-dependence
The same words can be toxic, or reclaimed, depending on speaker and context. This is a well-documented problem, specifically for slur reclamation and in-group language. Sap et al. (2019) show annotator disagreement on toxicity is itself correlated with annotator identity and dialect familiarity. Even "ground truth" labels for this task carry real, measurable subjectivity, not only noise.
Evaluation beyond accuracy
Standard per-label precision, recall and F1 apply. Most real-world text is non-toxic, and subgroup bias is well documented. So evaluation increasingly reports subgroup AUC alongside overall metrics. This catches the case where overall accuracy looks fine while one identity subgroup is systematically mis-scored.
Key references
- cjadams et al. (2017). Toxic Comment Classification Challenge. Kaggle / Jigsaw.
- Dixon, L. et al. (2018). Measuring and Mitigating Unintended Bias in Text Classification. AIES.
- Borkan, D. et al. (2019). Nuanced Metrics for Measuring Unintended Bias. arXiv:1903.04561
- Sap, M. et al. (2019). The Risk of Racial Bias in Hate Speech Detection. ACL.
Current state and open problems
LLM-based moderation can be prompted with a platform's specific policy text, rather than a fixed label set. It increasingly supplements or replaces fixed-taxonomy classifiers, offering more flexible, context-sensitive judgement at higher per-item compute cost.
The open problem remains the one Dixon et al. named in 2018: distinguishing genuine toxicity from surface-level correlation with identity terms, without shutting down legitimate discussion of those identities. No current method eliminates this trade-off entirely. Production systems manage it through ongoing subgroup evaluation and human review, not one single fix.
What to learn next
- Bias in datasets — the general problem this lesson's identity-term bias is one instance of.
- Fairness metrics — the formal metrics behind subgroup evaluation.
- When users write 'f r e e m0ney' to dodge your filter — the adversarial evasion problem in full.