The problem

Keyword filters punished the wrong players.

In a game, toxicity is not a nuisance, it is churn. Aggression, racism and bullying in the chat rooms were driving away exactly the new and paying players the platform needed, because they did not feel safe enough to stay.

The tools meant to stop it were blunt. They blocked on keywords, so they missed disguised abuse and the fast-moving slang of gamer culture, while flagging perfectly ordinary phrases as violations. Honest players got muted for nothing, which made them just as angry as the toxicity did.

And the legal pressure was real. Child-protection and online-safety regulation was tightening across Europe and the US, demanding strict enforcement and fast reporting, and the company was struggling to meet rules that differed country by country. Its licences were at risk.

What we did

Understand the sentence, not just the word.

A language engine that reads context and culture, enforcement that fits the offence, and rules that follow the law.

Context-aware

Slang in dozens of languages

We built a language engine trained on community slang and gamer speech across more than 30 languages and dialects. It reads the whole sentence and its context, so it can tell competitive trash-talk from genuine harassment.

Graduated response

The punishment fits the offence

Enforcement escalates in real time: an automatic warning, a temporary mute, or a flag for fast human review, so a first slip is handled differently from sustained abuse, and the rest of the room is never disrupted.

Fewer false blocks

Honest players stop getting muted

Because the model understands intent, it stopped punishing legitimate speech and local expressions that the old keyword filters had wrongly blocked, which kept the flow and freedom fair players expect.

Geo-aware

Rules that follow the border

Filtering rules change dynamically with the player's region, so the platform meets each country's child-protection and online-safety law precisely, rather than applying one blunt global policy.

Voice and text

Every channel watched

The engine covers both voice and text chat, so abuse cannot simply move to whichever channel is unmonitored.

Human-in-the-loop

People on the hard calls

Complex or borderline cases route to a human moderator with the offending moment flagged, so the difficult judgements stay with a person, made quickly.

The result

Safer rooms, fairer enforcement, licences secure.

The toxic players left, and the honest ones stayed.

Live

Toxicity down 45%, retention up 20%

Reports of toxic language and harassment in the chat rooms fell 45%, and retention of new players rose 20%, because the people the platform most wanted to keep finally felt safe enough to stay.

And fairer

False blocks down 75%, fully compliant

Wrongful blocks and false alarms dropped 75%, keeping the flow and freedom fair players expect, and the platform reached full compliance with local child-protection and online-safety regulation, lifting the threat of fines and shutdowns.

Why it holds

Moderation that punishes the innocent is not moderation.

The hard part of community safety is not catching slurs; it is telling banter from abuse, across languages and cultures, without muting the people you are trying to protect. That takes a model that reads context, enforcement that fits the offence, and rules that follow the local law. It is the same trust-and-safety discipline behind our marketplace integrity work, tuned for live chat instead of reviews.

More case studies

Related work.

Toxicity driving your community away?

Book a strategy call Bring the chat your filters get wrong both ways. Thirty minutes, no slides, or see more case studies.