AI Usage Quotas
Limit how much AI work a plan or workspace can do, and show what is left.
What it adds
Countable allowances for AI messages, runs, documents, or tokens, enforced per plan and workspace with a defined reset period.
What your agent is told to do
5
What your agent is told to do
5-
1
Identify the units the product actually sells on: messages, runs, processed documents, or tokens. Pick as few as possible and define each one precisely before writing any enforcement.
-
2
Enforce every quota on the server at the point the work is dispatched. Client-side counters and hidden buttons are presentation, not enforcement.
-
3
Decide and document whether a run that failed, was cancelled, or returned a refusal consumes allowance, and apply that decision consistently across every AI surface.
-
4
Support a shared workspace pool with optional per-member caps, so one member cannot drain the whole allowance in an afternoon.
-
5
AI Cost Budgets owns money and provider spend. This entry owns countable units and plan entitlements. Reuse the same dispatch interception point but do not derive one limit from the other.
Edge cases it handles
8
Edge cases it handles
8- Any check performed only in the interface can be bypassed by calling the endpoint directly. The dispatch path must re-check the quota server-side every time, including for background and scheduled runs.
- Failed, cancelled, and refused runs need an explicit rule. Charging for a run the model refused feels like theft; charging for nothing at all invites retry loops as a way around the limit.
- A shared workspace pool without a per-member cap lets one person exhaust the month for everyone. Support both, and make the interaction between them explicit in the interface.
- Reset periods must be anchored to a stated time zone and to the billing anniversary rather than a server-local midnight, or usage will reset at a time that makes no sense to the customer.
- Token counts are estimates until the provider reports actuals. Show remaining usage in the honest unit and never present an estimate with a precision it does not have.
- Concurrent requests near the limit must not all be admitted. Decrement atomically at dispatch rather than reading a count and writing it back.
- A plan downgrade mid-period must not produce a negative balance or retroactively invalidate work already done.
- When a quota is exhausted, the user needs to know which unit ran out, when it resets, and who in the workspace can raise it. A generic refusal produces a support ticket.
Definition of done
9
Definition of done
9- Quotas are enforced server-side on every dispatch path, including background and scheduled runs.
- The treatment of failed, cancelled, and refused runs is defined and applied uniformly.
- Workspace pools and per-member caps both exist and their interaction is visible to the user.
- Reset periods use a stated time zone and align to the billing anniversary.
- Remaining usage is displayed in honest units with no false precision on estimates.
- Concurrent requests at the limit cannot exceed the allowance.
- An exhausted quota tells the user what ran out, when it resets, and who can raise it.
- The feature matches the existing design system.
- No existing functionality is broken.
Related features
AI Cost Budgets
AI Cost Budgets
Cap what AI features are allowed to spend before the bill arrives.
What it does
Monetary spending limits on AI work, scoped by workspace, feature, and time period, enforced before a run starts.
How it works
- 1 Find every place the app calls a model and route all of them through one accounting point that records estimated and actual spend against a scope. A budget that only covers the chat feature is not a budget.
- 2 Estimate the cost of a run from the size of its input before dispatching it, and refuse anything that would exceed the remaining budget on its own.
- 3 Reserve the estimate against the budget when the run starts, then reconcile to the real usage figures when it finishes, releasing whatever was over-reserved.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/ai-cost-budgets
Multi-Model Routing
Multi-Model Routing
Send each AI request to the right model using rules you can read and test.
What it does
A deterministic routing layer that picks a model per request from task type, context size, latency budget, and data sensitivity.
How it works
- 1 Express routing as explicit, ordered rules over inputs the app can measure: task type, estimated context size, latency budget, and the sensitivity classification of the data involved. A rule set that can be read line by line can be reviewed and tested.
- 2 Make routing deterministic. The same inputs must always produce the same route, so a bad output can be reproduced and a rule change can be evaluated. Randomised or load-based selection turns every incident into guesswork.
- 3 Classify data before routing and refuse to route restricted content to any destination not approved for it. This check is a hard block, not a preference, and it must run before the request is assembled.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/multi-model-routing
Model Selection
Model Selection
Let each AI task run on the model that suits its quality, speed, and cost needs.
What it does
A per-task model choice, drawn from the models the app already has configured, with capability filtering and safe defaults.
How it works
- 1 Enumerate the models the app already has access to and record what each one can actually do: context capacity, whether it can return the structured output the task requires, whether it supports the tools the task calls, and its relative cost and speed.
- 2 Offer only the models that satisfy the task's requirements. A task that needs structured output must not list a model that cannot reliably produce it, because the failure appears later as malformed responses rather than as an unavailable option.
- 3 Store the choice against the specific task, not as one global setting. A single default forces a summarisation task and a classification task onto the same tier when they have opposite needs.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/model-selection
How it works
-
1
Copy the link
Grab the Markdown instruction URL for this feature.
-
2
Give it to your AI
Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.
-
3
It inspects, then implements
Your agent reads your existing app first, then adds the feature to fit it.
Works with your stack
These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.
Need it tighter than that? Customize the feature and tell it exactly what you're running.