Raju Shingadiya Raju Shingadiya
All posts

A toxicity filter playbook: 6 places to put Gatekeeper in your app

A toxicity model on its own does nothing. What matters is where you check, how strict you are there, and what happens next. Here is a playbook of six places, with the setting and code for each.

Gatekeeper desktop sample checking a message for toxicity
One call per message, about a millisecond, nothing sent anywhere.

Before the plays: the three settings

Gatekeeper reads a message and returns a score from 0 to 1, a yes/no verdict and a hint about the category: insult, profanity, threat, hate or sexual. The verdict depends on how strict you ask it to be:

SensitivityThresholdToxic caughtNormal chat flaggedGood for
STRICT0.25~85%~2–3%kids' apps, sending to a human
BALANCED0.42~77%~1%most apps (default)
RELAXED0.70~64%~0.6%hiding things automatically

The rule of thumb that falls out of this table: be strict when a human decides next, relaxed when the machine decides alone. Every play below follows it.

// build.gradle.kts
implementation("io.github.rajumark:gatekeeper:2.0.0")

val gatekeeper = Gatekeeper()   // load once (~50 ms), off the main thread; check() is thread-safe

The six plays

1. The nudge before sending

BALANCEDon the sender's deviceuser decides

Check the draft when the user taps send. If it's toxic, don't block — ask: “This might come across as hurtful. Send anyway?” The goal is a second thought, not censorship.

if (gatekeeper.isToxic(draft)) showRethinkDialog(draft) else send(draft)

2. Blur on arrival

RELAXEDon the receiver's devicemachine decides

Incoming messages that score high get blurred with a “tap to show”. The receiver stays in control, and because it runs on their device, it works even in end-to-end encrypted chats where a server can't read anything.

val hidden = gatekeeper.check(message.text, Sensitivity.RELAXED).isToxic

3. The moderator queue

STRICTcomments, reviews, postshuman decides

For public content, flag generously and let a person look. Sort the queue by score so the worst is seen first. A 2–3% false-flag rate is fine when the cost of a false flag is one extra glance.

val v = gatekeeper.check(comment, Sensitivity.STRICT)
if (v.isToxic) moderationQueue.add(comment, priority = v.score)

4. Kids' mode

STRICTyoung audiencesblock + tell a parent

With children, a missed insult is worse than a blocked joke. Use STRICT, and for threats specifically, don't just hide the message — surface it to whoever is responsible.

val v = gatekeeper.check(text, Sensitivity.STRICT)
if (v.isToxic && v.topCategory == Category.THREAT) notifyGuardian(text)

5. Names and bios

BALANCEDusernames, group names, profilesask to change

A toxic group name is seen by everyone in the group. Check names and bios at the moment they're saved and ask for another one. It's short text, so it costs almost nothing.

6. Your own custom line

custom thresholdtuned on your data

If none of the three presets fit, pass your own threshold. Take a few hundred real messages from your app, look at their scores, and pick the line where you're comfortable.

gatekeeper.check(message, threshold = 0.6f)
gatekeeper.score(message)   // just the number, 0..1

Why it's built for Indian chat

Most toxicity models are trained on English, and a little on other languages written in their own scripts. Indian chat is Hindi typed in Latin letters, mixed with English, with creative spelling. That's where a generic model falls over, and where Gatekeeper was trained to hold up:

Gatekeeper letting a friendly Hinglish message through
“bhai tu toh kamaal hai ๐Ÿ”ฅ” → fine
Gatekeeper catching a threat written in Hinglish
“tujhe jaan se maar dunga” → threat
Gatekeeper catching an English insult
“you are a stupid idiot” → insult
GatekeeperMultilingual BERT toxicity classifier
Indian-language comments (hi, ta, te, ml, kn), F10.840.63
Same comments in Latin letters, F10.830.65
Everyday chat not flagged98.8%85.1%
15 world languages, F10.770.91
Hate-speech stress test (no slurs)45%61%
Size3.8 MB711 MB

The 98.8% row matters more than it looks. A filter that flags one in seven normal messages (85.1%) trains users to ignore it. Slang like “this movie killed me ๐Ÿ˜‚” should stay clean, and it does.

What not to do

The short version

Strict where a human reviews. Relaxed where the app acts alone. Balanced for everything in between. Run it on the device so it works in private chats too, and keep a person in the loop for anything that really matters.


The Hoverfly series. Eight small on-device models, one job each, all plain Kotlin Multiplatform with no native code. Each post is written differently, because each model taught me something different.

  • Moji — emoji suggestions, told as a story
  • Beacon — language detection, as a benchmark report
  • Comma — punctuation restoration, as a step-by-step tutorial
  • Comeback — smart replies, as an honest opinion piece
  • Emotion — emotion detection, as questions and answers
  • Gatekeeper — toxicity detection, as a practical playbook (you are here)
  • Hideout — personal info hiding, as an audit checklist
  • Chalk — doodle recognition, as an engineering deep dive
← All posts