Vector Index Health
Show operators whether each knowledge source is complete, delayed, stale, or failing.
What it adds
An operator view of index state per source: expected versus indexed counts, embedding age, queue depth, and failures.
What your agent is told to do
5
What your agent is told to do
5-
1
Record the expected document and chunk counts for every source alongside what is actually indexed, so a shortfall shows as a number on a screen rather than being discovered through a missing answer.
-
2
Stamp every stored embedding with the content version and the embedding configuration that produced it, and mark it stale when either changes. Without that stamp there is no way to tell current vectors from ones generated under an older setup.
-
3
Present queue delay and permanent failure as different states with different actions. A backlog clears itself and needs patience; a failing source needs a person, and merging the two hides real breakage behind a spinner.
-
4
Offer a re-index that builds the new vectors and swaps them in for a source, retiring the old set as part of the same operation. Do not re-index by writing fresh vectors and leaving the previous ones in place, because duplicates skew retrieval scoring.
-
5
Restrict this view to operators. It exposes document titles, counts, and raw error text across tenants. Per-answer inspection of which chunks produced a specific result belongs to Retrieval Debugger.
Edge cases it handles
8
Edge cases it handles
8- Expected and indexed counts must be compared at both the document and the chunk level, since a document can be present but only partly chunked.
- Changing the embedding configuration invalidates every existing vector at once. Flag the whole affected set as stale and show the scale of the re-index before anyone starts it.
- A queue that is merely deep looks identical to one that is stuck. Show the age of the oldest waiting item and the rate of progress so the difference is visible.
- Re-indexing must be safe to run twice. A second run started because the first looked slow must converge on one set of vectors, not two.
- Error text from an embedding provider or an external source can contain customer content, so the whole surface must be behind operator permission and excluded from general logging.
- A full re-index costs money and consumes provider rate limit. Show an estimate before it runs, make it interruptible, and keep the existing index serving queries throughout.
- A source deleted while a sync is running will leave counts that never reconcile. Reconcile against the source's current state rather than the state at the time the run started.
- A source with genuinely zero documents is healthy, not broken, and must read differently from a source whose ingestion produced nothing.
Definition of done
8
Definition of done
8- Every source shows expected versus indexed document and chunk counts.
- Embeddings carry the content version and configuration they were produced under, and stale ones are flagged.
- Queue delay, in-progress work, and permanent failure are distinct states with distinct actions.
- Re-indexing a source replaces its vectors without leaving duplicates, and is safe to run more than once.
- Queries keep being served from the existing index while a re-index runs.
- The view and its error text are restricted to authorized operators.
- The feature matches the existing design system.
- No existing functionality is broken.
Related features
AI Cost Budgets
AI Cost Budgets
Cap what AI features are allowed to spend before the bill arrives.
What it does
Monetary spending limits on AI work, scoped by workspace, feature, and time period, enforced before a run starts.
How it works
- 1 Find every place the app calls a model and route all of them through one accounting point that records estimated and actual spend against a scope. A budget that only covers the chat feature is not a budget.
- 2 Estimate the cost of a run from the size of its input before dispatching it, and refuse anything that would exceed the remaining budget on its own.
- 3 Reserve the estimate against the budget when the run starts, then reconcile to the real usage figures when it finishes, releasing whatever was over-reserved.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/ai-cost-budgets
Multi-Model Routing
Multi-Model Routing
Send each AI request to the right model using rules you can read and test.
What it does
A deterministic routing layer that picks a model per request from task type, context size, latency budget, and data sensitivity.
How it works
- 1 Express routing as explicit, ordered rules over inputs the app can measure: task type, estimated context size, latency budget, and the sensitivity classification of the data involved. A rule set that can be read line by line can be reviewed and tested.
- 2 Make routing deterministic. The same inputs must always produce the same route, so a bad output can be reproduced and a rule change can be evaluated. Randomised or load-based selection turns every incident into guesswork.
- 3 Classify data before routing and refuse to route restricted content to any destination not approved for it. This check is a hard block, not a preference, and it must run before the request is assembled.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/multi-model-routing
Model Selection
Model Selection
Let each AI task run on the model that suits its quality, speed, and cost needs.
What it does
A per-task model choice, drawn from the models the app already has configured, with capability filtering and safe defaults.
How it works
- 1 Enumerate the models the app already has access to and record what each one can actually do: context capacity, whether it can return the structured output the task requires, whether it supports the tools the task calls, and its relative cost and speed.
- 2 Offer only the models that satisfy the task's requirements. A task that needs structured output must not list a model that cannot reliably produce it, because the failure appears later as malformed responses rather than as an unavailable option.
- 3 Store the choice against the specific task, not as one global setting. A single default forces a summarisation task and a classification task onto the same tier when they have opposite needs.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/model-selection
How it works
-
1
Copy the link
Grab the Markdown instruction URL for this feature.
-
2
Give it to your AI
Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.
-
3
It inspects, then implements
Your agent reads your existing app first, then adds the feature to fit it.
Works with your stack
These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.
Need it tighter than that? Customize the feature and tell it exactly what you're running.