Retrieval Debugger
Show exactly which sources, chunks, and scores produced a given AI answer.
What it adds
A per-answer inspector showing the query as issued, the filters applied, the candidate chunks with their scores, and what reached the model.
What your agent is told to do
5
What your agent is told to do
5-
1
Capture for each answer the query as it was issued, the filters applied, the candidates returned with their scores, and which of those actually made it into the request after the context ceiling was applied.
-
2
Show results after permission filtering, with a count of how many candidates were excluded and why. Displaying the pre-filter set turns the debugger into a way to read content the viewer cannot open.
-
3
Present each scoring stage separately — keyword, semantic, and any reranking — because a chunk that ends up first overall may have been rescued by one stage after being buried by another, and a single blended number hides that.
-
4
Support replaying a captured query against the recorded index version, and label the replay clearly when the index has moved on since, so a comparison is never mistaken for the original run.
-
5
Do not attach debug metadata to normal end-user responses. Keep the whole surface behind operator permission, and leave index-wide counts and staleness to Vector Index Health rather than repeating them per answer.
Edge cases it handles
8
Edge cases it handles
8- Retrieval must be shown as the user experienced it, meaning after tenant and record filtering. A debugger that reports the unfiltered candidate set is reporting something that never influenced the answer.
- Chunk text the operator viewing the debugger has no right to read must be redacted to metadata — source, score, length — rather than displayed in full.
- Scores from different stages are not on the same scale and must not be plotted as though they were. Label each and show the ranking each stage produced.
- Replay needs the query, the filters, and the index version pinned together. Replaying against a changed index produces a different answer for reasons unrelated to the change being investigated.
- Debug records must never leak into ordinary responses, logs shown to customers, or error messages, since they contain retrieved content verbatim.
- Debug records hold customer content and therefore need their own retention window and deletion, shorter than the records they describe.
- Some answers are produced with no retrieval at all, and some with retrieval that returned nothing. Both need a clear representation rather than an empty table that looks like a capture failure.
- Replaying a query calls the model again and costs money, so it must be an explicit action with a visible warning, never something that happens on page load.
Definition of done
9
Definition of done
9- Every answer has a stored capture of its query, filters, candidates, scores, and final context.
- Results are displayed after permission filtering, with excluded candidates counted rather than shown.
- Content the viewing operator cannot access is redacted to metadata.
- Keyword, semantic, and rerank scores are shown separately and labelled.
- A query can be replayed against its recorded index version, with drift clearly marked.
- No debug metadata appears in end-user responses.
- Debug captures have their own retention period and are deleted when it expires.
- The feature matches the existing design system.
- No existing functionality is broken.
Related features
AI Cost Budgets
AI Cost Budgets
Cap what AI features are allowed to spend before the bill arrives.
What it does
Monetary spending limits on AI work, scoped by workspace, feature, and time period, enforced before a run starts.
How it works
- 1 Find every place the app calls a model and route all of them through one accounting point that records estimated and actual spend against a scope. A budget that only covers the chat feature is not a budget.
- 2 Estimate the cost of a run from the size of its input before dispatching it, and refuse anything that would exceed the remaining budget on its own.
- 3 Reserve the estimate against the budget when the run starts, then reconcile to the real usage figures when it finishes, releasing whatever was over-reserved.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/ai-cost-budgets
Multi-Model Routing
Multi-Model Routing
Send each AI request to the right model using rules you can read and test.
What it does
A deterministic routing layer that picks a model per request from task type, context size, latency budget, and data sensitivity.
How it works
- 1 Express routing as explicit, ordered rules over inputs the app can measure: task type, estimated context size, latency budget, and the sensitivity classification of the data involved. A rule set that can be read line by line can be reviewed and tested.
- 2 Make routing deterministic. The same inputs must always produce the same route, so a bad output can be reproduced and a rule change can be evaluated. Randomised or load-based selection turns every incident into guesswork.
- 3 Classify data before routing and refuse to route restricted content to any destination not approved for it. This check is a hard block, not a preference, and it must run before the request is assembled.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/multi-model-routing
Model Selection
Model Selection
Let each AI task run on the model that suits its quality, speed, and cost needs.
What it does
A per-task model choice, drawn from the models the app already has configured, with capability filtering and safe defaults.
How it works
- 1 Enumerate the models the app already has access to and record what each one can actually do: context capacity, whether it can return the structured output the task requires, whether it supports the tools the task calls, and its relative cost and speed.
- 2 Offer only the models that satisfy the task's requirements. A task that needs structured output must not list a model that cannot reliably produce it, because the failure appears later as malformed responses rather than as an unavailable option.
- 3 Store the choice against the specific task, not as one global setting. A single default forces a summarisation task and a classification task onto the same tier when they have opposite needs.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/model-selection
How it works
-
1
Copy the link
Grab the Markdown instruction URL for this feature.
-
2
Give it to your AI
Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.
-
3
It inspects, then implements
Your agent reads your existing app first, then adds the feature to fit it.
Works with your stack
These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.
Need it tighter than that? Customize the feature and tell it exactly what you're running.