AddThisFeature

Prompt Injection Defense

Stop instructions hidden in documents, pages, and tool output from steering the AI.

involved AI Safety & Trust

What it adds

A boundary between the app's own instructions and untrusted content, backed by enforced tool permissions and confirmation for consequential actions.

What your agent is told to do

5
  1. 1

    Enumerate every route by which content the app did not author reaches a model: uploaded files, retrieved passages, fetched pages, records from connected accounts, and the results of tool calls.

  2. 2

    Keep the app's own rules in the trusted portion of the request and enclose untrusted content in a clearly delimited region labelled as material to be analysed rather than instructions to be obeyed.

  3. 3

    Enforce what the model may reach outside the model. Tool permissions, tenant scoping, and rate limits must hold even when the model has been completely persuaded — an instruction telling the model not to delete things is not a control.

  4. 4

    Require explicit human confirmation for any consequential action — sending, paying, deleting, sharing, or changing permissions — whenever untrusted content was in the context that produced it, and show the user which content prompted it.

  5. 5

    Do not rely on a detector as the defense. Pattern matching catches the obvious attempts and misses the rest; use it to flag and record suspicious content while the permission boundary remains the thing that actually prevents harm.

Edge cases it handles

8
  • Retrieved passages and tool results are data. Anything in them that reads as a command — including text claiming to come from the system or the developer — must be treated as content under analysis, never as an instruction to follow.
  • The trusted rules must be structurally separate from untrusted context, so that content ending a delimiter early, or imitating the app's own framing, cannot dissolve the boundary.
  • Tool access and data scope must be restricted by the calling user's permissions independently of what the model asks for, so a persuaded model simply gets a denial.
  • Detection is best effort by definition. Log what it flags, review it, and never let a clean detection result be the reason an action proceeds unconfirmed.
  • Any action influenced by external content needs a confirmation step that names what will happen and to what, because a user approving a vague summary is not really approving anything.
  • Injections hide in places nobody inspects: alt text, filenames, spreadsheet cells, document metadata, and text rendered invisibly in a fetched page. Normalize and strip content before it goes anywhere near a request.
  • Exfiltration usually comes disguised as output rather than as an action. Treat model-generated links, image addresses, and outbound requests as untrusted and do not fetch or render them automatically.
  • A compromised run can report success while doing nothing or doing something else. Verify the outcome from the application's own state rather than from the model's account of what it did.

Definition of done

9
  • Every path by which external content reaches a model is identified and passes through the same untrusted-content handling.
  • Application rules and untrusted content are structurally separated in every request.
  • Tool and data access are enforced by the application against the user's permissions, independently of model output.
  • Consequential actions influenced by external content require an explicit confirmation naming the action and its target.
  • Suspicious content is flagged and recorded, and detection alone never authorizes an action.
  • Model-generated links and resource addresses are not fetched or rendered automatically.
  • Content in metadata, filenames, and hidden markup is normalized before it enters a request.
  • The feature matches the existing design system.
  • No existing functionality is broken.

Related features

How it works

  1. 1

    Copy the link

    Grab the Markdown instruction URL for this feature.

  2. 2

    Give it to your AI

    Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.

  3. 3

    It inspects, then implements

    Your agent reads your existing app first, then adds the feature to fit it.

Works with your stack

These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.

Need it tighter than that? Customize the feature and tell it exactly what you're running.