AI Image Content Moderation
Check uploaded images against the app's content rules before anyone else sees them.
What it adds
An automated classification step on image upload that holds, flags, or clears content against per-category policy thresholds.
What your agent is told to do
5
What your agent is told to do
5-
1
Find every path by which an image becomes visible to someone other than its uploader, including avatars, attachments, embeds, and public links, and put the check on all of them. A single unguarded upload path defeats the feature.
-
2
Express policy as named categories with their own thresholds, each mapped to an action. A single unsafe score gives operators no way to be strict about one category and permissive about another.
-
3
Hold anything above the review threshold in a state where the uploader can still see their own file but nobody else can, and reuse the moderation queue, action vocabulary, and appeal flow shared with AI Text Content Moderation rather than building a second one.
-
4
Record every decision with the image reference, the categories and scores, the action taken, and who or what took it, and keep that record long enough to handle an appeal. Do not retain the image itself beyond the app's normal storage rules for the sake of the audit trail.
-
5
Do not surface category names, scores, or thresholds to the uploader. Tell them the content was held for review and how to appeal; publishing the classifier's reasoning teaches people how to get around it.
Edge cases it handles
8
Edge cases it handles
8- Per-category thresholds tuned to the app's actual policy are required. One generic score forces a choice between blocking harmless content and admitting content the policy forbids.
- Scores in the uncertain band must go to a human queue rather than defaulting to allow or to block, and that queue needs a service-level expectation so held content does not sit for a week.
- Content held for review must be genuinely inaccessible to everyone except the uploader and reviewers, including through direct storage links, cached copies, and thumbnails already generated.
- False positives are certain. Give the uploader a stated appeal path, let a reviewer overturn a decision, and make the overturn restore visibility without requiring a re-upload.
- Do not expose category names, confidence scores, or thresholds in responses, client code, or error messages, because each one is a hint for evading the check.
- When the moderation provider is unavailable or times out, choose a documented failure posture — hold for review rather than publish — and make sure held items are re-checked automatically when service returns.
- Rate limits and large batches need backoff and a queue on the app's existing background-job system, with the uploader told their image is pending rather than shown an error.
- Re-uploading the same file must reach the same decision without a second classification, and a policy change must be able to trigger a re-check of previously cleared content.
Definition of done
9
Definition of done
9- Every path that publishes an image runs the check before the image is visible to others.
- Policy is expressed as named categories with independently configurable thresholds and actions.
- Held images are inaccessible to everyone but the uploader and reviewers, including through direct storage and thumbnail URLs.
- Uncertain results reach a human review queue rather than defaulting to allow or block.
- Uploaders have an appeal path and an overturned decision restores visibility without a re-upload.
- Category names, scores, and thresholds are never returned to the uploader.
- A provider outage holds content for review and re-checks it automatically when service returns.
- The feature matches the existing design system.
- No existing functionality is broken.
Related features
Prompt Injection Defense
Prompt Injection Defense
Stop instructions hidden in documents, pages, and tool output from steering the AI.
What it does
A boundary between the app's own instructions and untrusted content, backed by enforced tool permissions and confirmation for consequential actions.
How it works
- 1 Enumerate every route by which content the app did not author reaches a model: uploaded files, retrieved passages, fetched pages, records from connected accounts, and the results of tool calls.
- 2 Keep the app's own rules in the trusted portion of the request and enclose untrusted content in a clearly delimited region labelled as material to be analysed rather than instructions to be obeyed.
- 3 Enforce what the model may reach outside the model. Tool permissions, tenant scoping, and rate limits must hold even when the model has been completely persuaded — an instruction telling the model not to delete things is not a control.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/prompt-injection-defense
AI Import Column Mapping
AI Import Column Mapping
Guess how an uploaded file's columns line up with the app's fields, then ask.
What it does
A proposed mapping from uploaded columns to destination fields, shown with samples and confirmed before the import runs.
How it works
- 1 Take the uploaded headers and a small sample of rows and propose a mapping to the destination fields. Match on exact and near-exact header names first and only consult the model for the columns that remain unresolved.
- 2 Constrain the candidate set to fields the current user and the importer are actually permitted to write. A mapping that targets a read-only, computed, or privileged field must never be offered, whatever the header says.
- 3 Show the proposal as a table: source column, sample values, destination field, confidence, and the transformation that will be applied. A user cannot approve a mapping they cannot see the effect of.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/ai-import-column-mapping
Model Fallback
Model Fallback
Keep AI features working when the primary model provider fails.
What it does
A bounded retry path that reruns a failed AI task on a compatible alternative model and discloses the substitution.
How it works
- 1 Classify failures before reacting. Provider outages, capacity rejections, rate limiting, and timeouts are worth retrying elsewhere; a refusal, an invalid request, or an authentication error will fail identically on any destination and must surface immediately.
- 2 Honour rate limiting with backoff before moving on. Retrying instantly against a provider that asked you to wait deepens the outage and can extend the penalty period.
- 3 Allow fallback only to models that can honour the same output shape and the same tool behaviour as the original. A substitute that returns a different structure turns a provider outage into a data problem inside the app.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/model-fallback
How it works
-
1
Copy the link
Grab the Markdown instruction URL for this feature.
-
2
Give it to your AI
Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.
-
3
It inspects, then implements
Your agent reads your existing app first, then adds the feature to fit it.
Works with your stack
These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.
Need it tighter than that? Customize the feature and tell it exactly what you're running.