SafetyAI
The text half of the scam problem. A cheap regex pass runs first; anything suspicious gets a few-shot LLM verdict. It learns from moderator feedback, so it gets better without getting jumpier.
[01] why
The right check in the right order
Fake Nitro, phishing links, giveaways that want your login. Keyword lists catch the lazy ones and miss the rest. Sending every message to an AI catches more, but it's slow and it costs money. SafetyAI does both, in that order.
[02] how it works
From regex to verdict
- A regex pre-filter runs first. If none of the trigger patterns match, the AI is never called.
- Hard-block patterns mark a message as a scam straight away.
- Anything suspicious goes to a language model with a few-shot prompt built from real examples.
- Moderators confirm or reject with a button, and that answer becomes a new example. That's how it learns without getting jumpier.
[03] safety switches
Built to be trusted
- productionReadydeleting stays off until you flip it
- confidencea minimum score before it deletes anything
- budgetglobal token bucket + per-channel cooldown
- actionsdelete, DM the user, mod alert with buttons
[04] providers
Bring your own model
- OpenAI
- Anthropic
- Mistral
- Groq
- xAI
- Hugging Face
Pick one in the config and give it a key. For images, there's Argus.
[05] built with
Made with
- TypeScript
- Node.js
- Pino