05open sourceanti-scam2025

SafetyAI

The text half of the scam problem. A cheap regex pass runs first; anything suspicious gets a few-shot LLM verdict. It learns from moderator feedback, so it gets better without getting jumpier.

[01] why

The right check in the right order

Fake Nitro, phishing links, giveaways that want your login. Keyword lists catch the lazy ones and miss the rest. Sending every message to an AI catches more, but it's slow and it costs money. SafetyAI does both, in that order.

[02] how it works

From regex to verdict

  • A regex pre-filter runs first. If none of the trigger patterns match, the AI is never called.
  • Hard-block patterns mark a message as a scam straight away.
  • Anything suspicious goes to a language model with a few-shot prompt built from real examples.
  • Moderators confirm or reject with a button, and that answer becomes a new example. That's how it learns without getting jumpier.

[03] safety switches

Built to be trusted

  • productionReadydeleting stays off until you flip it
  • confidencea minimum score before it deletes anything
  • budgetglobal token bucket + per-channel cooldown
  • actionsdelete, DM the user, mod alert with buttons

[04] providers

Bring your own model

  • OpenAI
  • Anthropic
  • Mistral
  • Groq
  • xAI
  • Hugging Face

Pick one in the config and give it a key. For images, there's Argus.

[05] built with

Made with

  • TypeScript
  • Node.js
  • Pino