Prompt Injection Defense
Stop instructions hidden in documents, pages, and tool output from steering the AI.
What it adds
A boundary between the app's own instructions and untrusted content, backed by enforced tool permissions and confirmation for consequential actions.
What your agent is told to do
5
What your agent is told to do
5-
1
Enumerate every route by which content the app did not author reaches a model: uploaded files, retrieved passages, fetched pages, records from connected accounts, and the results of tool calls.
-
2
Keep the app's own rules in the trusted portion of the request and enclose untrusted content in a clearly delimited region labelled as material to be analysed rather than instructions to be obeyed.
-
3
Enforce what the model may reach outside the model. Tool permissions, tenant scoping, and rate limits must hold even when the model has been completely persuaded — an instruction telling the model not to delete things is not a control.
-
4
Require explicit human confirmation for any consequential action — sending, paying, deleting, sharing, or changing permissions — whenever untrusted content was in the context that produced it, and show the user which content prompted it.
-
5
Do not rely on a detector as the defense. Pattern matching catches the obvious attempts and misses the rest; use it to flag and record suspicious content while the permission boundary remains the thing that actually prevents harm.
Edge cases it handles
8
Edge cases it handles
8- Retrieved passages and tool results are data. Anything in them that reads as a command — including text claiming to come from the system or the developer — must be treated as content under analysis, never as an instruction to follow.
- The trusted rules must be structurally separate from untrusted context, so that content ending a delimiter early, or imitating the app's own framing, cannot dissolve the boundary.
- Tool access and data scope must be restricted by the calling user's permissions independently of what the model asks for, so a persuaded model simply gets a denial.
- Detection is best effort by definition. Log what it flags, review it, and never let a clean detection result be the reason an action proceeds unconfirmed.
- Any action influenced by external content needs a confirmation step that names what will happen and to what, because a user approving a vague summary is not really approving anything.
- Injections hide in places nobody inspects: alt text, filenames, spreadsheet cells, document metadata, and text rendered invisibly in a fetched page. Normalize and strip content before it goes anywhere near a request.
- Exfiltration usually comes disguised as output rather than as an action. Treat model-generated links, image addresses, and outbound requests as untrusted and do not fetch or render them automatically.
- A compromised run can report success while doing nothing or doing something else. Verify the outcome from the application's own state rather than from the model's account of what it did.
Definition of done
9
Definition of done
9- Every path by which external content reaches a model is identified and passes through the same untrusted-content handling.
- Application rules and untrusted content are structurally separated in every request.
- Tool and data access are enforced by the application against the user's permissions, independently of model output.
- Consequential actions influenced by external content require an explicit confirmation naming the action and its target.
- Suspicious content is flagged and recorded, and detection alone never authorizes an action.
- Model-generated links and resource addresses are not fetched or rendered automatically.
- Content in metadata, filenames, and hidden markup is normalized before it enters a request.
- The feature matches the existing design system.
- No existing functionality is broken.
Related features
AI Import Column Mapping
AI Import Column Mapping
Guess how an uploaded file's columns line up with the app's fields, then ask.
What it does
A proposed mapping from uploaded columns to destination fields, shown with samples and confirmed before the import runs.
How it works
- 1 Take the uploaded headers and a small sample of rows and propose a mapping to the destination fields. Match on exact and near-exact header names first and only consult the model for the columns that remain unresolved.
- 2 Constrain the candidate set to fields the current user and the importer are actually permitted to write. A mapping that targets a read-only, computed, or privileged field must never be offered, whatever the header says.
- 3 Show the proposal as a table: source column, sample values, destination field, confidence, and the transformation that will be applied. A user cannot approve a mapping they cannot see the effect of.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/ai-import-column-mapping
Model Fallback
Model Fallback
Keep AI features working when the primary model provider fails.
What it does
A bounded retry path that reruns a failed AI task on a compatible alternative model and discloses the substitution.
How it works
- 1 Classify failures before reacting. Provider outages, capacity rejections, rate limiting, and timeouts are worth retrying elsewhere; a refusal, an invalid request, or an authentication error will fail identically on any destination and must surface immediately.
- 2 Honour rate limiting with backoff before moving on. Retrying instantly against a provider that asked you to wait deepens the outage and can extend the penalty period.
- 3 Allow fallback only to models that can honour the same output shape and the same tool behaviour as the original. A substitute that returns a different structure turns a provider outage into a data problem inside the app.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/model-fallback
AI Audit Trail
AI Audit Trail
Keep a reviewable record of what the AI saw, produced, recommended, and actually did.
What it does
A per-run record of inputs, sources, configuration, output, approvals, and executed actions, restricted and retained by purpose.
How it works
- 1 Write one record per run covering the feature invoked, the configuration and prompt version in force, the sources retrieved, the tools called, the output produced, who approved it, and what was carried out.
- 2 Distinguish a draft from an executed action explicitly in the record. A trail that shows only what text was produced cannot answer the question that matters in a review, which is whether anything happened.
- 3 Redact at the moment of writing. Strip credentials, tokens, and payment details before the record is stored, and keep only the personal content genuinely needed to understand the decision.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/ai-audit-trail
How it works
-
1
Copy the link
Grab the Markdown instruction URL for this feature.
-
2
Give it to your AI
Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.
-
3
It inspects, then implements
Your agent reads your existing app first, then adds the feature to fit it.
Works with your stack
These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.
Need it tighter than that? Customize the feature and tell it exactly what you're running.