Weapons and violence in photos — new moderation categories, included in your plan
From today modwall assesses images in three categories: nudity/NSFW (as before), weapons and violence/gore (beta). All included in your plan — one request, one limit, zero counting of "operations" per model.
Numbers, not promises
The weapons category (pistols, rifles, knives held in hand) was trained and evaluated on over 15k images, including "tricky" negatives — smartphones, wallets and banknotes held in hand like a weapon. Results on the validation set:
- AUC 0.99 — class separation on par with specialised detectors;
- at the default blocking threshold of 0.8: 99% precision — about 7 false alarms per 1000 ordinary photos — at a recall of 79%;
- with a review band of 0.5–0.8, recall rises to ~98% — borderline cases go to your human-in-the-loop queue rather than onto the platform.
Violence/gore currently runs in zero-shot mode (beta) — it returns a sensible signal, but calibrate the thresholds on your own traffic before enabling automatic blocking.
You decide what counts as a problem — per service
A weapon in a photo means one thing on a classifieds portal, another in a gamer community, and something else again on a platform for schools. That is why you enable categories in moderation profiles: you pin a profile to an API key or widget, and each of your services gets its own set of categories, thresholds and operating mode. A music service can turn off weapons (rappers…), a service for children can crank everything up to maximum.
How to enable it (2 minutes)
- Panel → Policy → choose a profile (or create a new one per service).
- Enable the Weapons and/or Violence category, set the thresholds.
- Start with monitor mode — you will see in the panel what would be flagged, with no changes for your users. When the numbers look good, switch to review or enforce.
The API response has gained the fields categories, blocked_by and
profile — details in the API reference. Nudity/NSFW works
as before, you don't need to change anything.
Why included in the plan?
At some providers each category is a separate model, and each model eats operations from the paid pool — moderating one image across three categories costs three operations. Our categories run on the same CPU infrastructure as the rest, so we see no reason to make you choose between safety and budget. One request = all enabled categories.
Create an account (1000 requests/mo for free) or test it on your own photo on the home page — the demo now shows all three categories.
Plug in moderation before a problem lands on your platform
1,000 free requests per month, no card. Monitor mode shows results on your real traffic — without blocking anything.
Create a free account Pricing