AI Cost Budgets
Cap what AI features are allowed to spend before the bill arrives.
What it adds
Monetary spending limits on AI work, scoped by workspace, feature, and time period, enforced before a run starts.
What your agent is told to do
5
What your agent is told to do
5-
1
Find every place the app calls a model and route all of them through one accounting point that records estimated and actual spend against a scope. A budget that only covers the chat feature is not a budget.
-
2
Estimate the cost of a run from the size of its input before dispatching it, and refuse anything that would exceed the remaining budget on its own.
-
3
Reserve the estimate against the budget when the run starts, then reconcile to the real usage figures when it finishes, releasing whatever was over-reserved.
-
4
Define both a soft limit that warns owners and continues, and a hard limit that stops new runs, and let an operator raise either without a deploy.
-
5
This entry owns money. AI Usage Quotas owns countable units such as messages, documents, and runs. Share the same interception point but keep the two ledgers separate, and do not express one in terms of the other.
Edge cases it handles
8
Edge cases it handles
8- An unusually large request must be priced before it is sent, not discovered afterwards. Estimate from the input size and refuse anything that alone would blow the remaining budget, telling the user what to trim.
- Streaming and tool-using runs do not have a known cost at dispatch. Reserve a conservative estimate up front, then reconcile against the real usage the provider reports when the run closes, releasing the difference.
- Several runs starting at the same moment must not each see the same remaining balance and all be admitted. Reserve atomically so concurrent requests cannot oversell the same headroom.
- Soft and hard limits must behave differently and visibly: a soft limit warns the owner and keeps working, a hard limit refuses new runs while letting in-flight ones finish and settle.
- The pricing figures used to compute spend change over time. Keep them versioned with effective dates, record which version priced each run, and never retroactively reprice settled history.
- If the provider never returns usage figures for a run, the reservation must still be settled on a timeout rather than pinning that amount forever.
- A budget reset at a period boundary must not release money for runs that are still in flight across it.
- When the model provider is unavailable, no budget is consumed. A failed dispatch must release its reservation rather than counting as spend.
Definition of done
8
Definition of done
8- Every model call in the app passes through a single accounting point that records spend against a scope.
- Large requests are cost-estimated before dispatch and refused with an explanation when they exceed the remaining budget.
- Reservations are made atomically and reconciled to actual usage, so concurrent runs cannot overspend a budget.
- Soft limits warn and continue while hard limits stop new work, and both are adjustable by an operator without a deploy.
- Pricing tables are versioned with effective dates and each run records the version that priced it.
- Failed, timed-out, and provider-unavailable runs release their reservations instead of counting as spend.
- The feature matches the existing design system.
- No existing functionality is broken.
Related features
Multi-Model Routing
Multi-Model Routing
Send each AI request to the right model using rules you can read and test.
What it does
A deterministic routing layer that picks a model per request from task type, context size, latency budget, and data sensitivity.
How it works
- 1 Express routing as explicit, ordered rules over inputs the app can measure: task type, estimated context size, latency budget, and the sensitivity classification of the data involved. A rule set that can be read line by line can be reviewed and tested.
- 2 Make routing deterministic. The same inputs must always produce the same route, so a bad output can be reproduced and a rule change can be evaluated. Randomised or load-based selection turns every incident into guesswork.
- 3 Classify data before routing and refuse to route restricted content to any destination not approved for it. This check is a hard block, not a preference, and it must run before the request is assembled.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/multi-model-routing
Model Selection
Model Selection
Let each AI task run on the model that suits its quality, speed, and cost needs.
What it does
A per-task model choice, drawn from the models the app already has configured, with capability filtering and safe defaults.
How it works
- 1 Enumerate the models the app already has access to and record what each one can actually do: context capacity, whether it can return the structured output the task requires, whether it supports the tools the task calls, and its relative cost and speed.
- 2 Offer only the models that satisfy the task's requirements. A task that needs structured output must not list a model that cannot reliably produce it, because the failure appears later as malformed responses rather than as an unavailable option.
- 3 Store the choice against the specific task, not as one global setting. A single default forces a summarisation task and a classification task onto the same tier when they have opposite needs.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/model-selection
Prompt Versioning
Prompt Versioning
Tie every AI output to the exact prompt version that produced it.
What it does
Immutable, numbered versions of each prompt, with the run configuration recorded and every output stamped with the version used.
How it works
- 1 Make every publish create a new immutable version rather than overwriting the previous text. Editing history in place destroys the only record of what produced last month's outputs.
- 2 Capture the whole run configuration with each version, not just the wording: which model tier and parameters were used, which tools were available, and the expected output shape. A prompt that behaves differently under different settings is not one prompt.
- 3 Stamp every generated output with the version identifier that produced it, and keep that stamp with the record so an output found later can be traced back to its exact instructions.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/prompt-versioning
How it works
-
1
Copy the link
Grab the Markdown instruction URL for this feature.
-
2
Give it to your AI
Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.
-
3
It inspects, then implements
Your agent reads your existing app first, then adds the feature to fit it.
Works with your stack
These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.
Need it tighter than that? Customize the feature and tell it exactly what you're running.