# AI Cost Budgets

## Objective

Cap what AI features are allowed to spend before the bill arrives.

Monetary spending limits on AI work, scoped by workspace, feature, and time period, enforced before a run starts.

## Before You Begin

This feature is being added to an application that already exists and already
works. Do not scaffold a new project, and do not assume a blank slate.

Inspect the codebase first and establish:

- The existing application structure and where code of this kind already lives.
- The framework and version in use.
- The existing design system — colours, spacing, typography, and component conventions.
- Existing UI components you can reuse instead of writing new ones.
- The existing database structure, if this feature needs to persist anything.
- The existing authentication and authorization system, if this feature is user-scoped.
- Dependencies already installed, so you don't add a library that duplicates one.
- The existing test setup and conventions.

Only start writing code once you understand the above. If the application
already implements part of this feature, extend it rather than replacing it.

## Implementation Instructions

1. Find every place the app calls a model and route all of them through one accounting point that records estimated and actual spend against a scope. A budget that only covers the chat feature is not a budget.
2. Estimate the cost of a run from the size of its input before dispatching it, and refuse anything that would exceed the remaining budget on its own.
3. Reserve the estimate against the budget when the run starts, then reconcile to the real usage figures when it finishes, releasing whatever was over-reserved.
4. Define both a soft limit that warns owners and continues, and a hard limit that stops new runs, and let an operator raise either without a deploy.
5. This entry owns money. AI Usage Quotas owns countable units such as messages, documents, and runs. Share the same interception point but keep the two ledgers separate, and do not express one in terms of the other.

## UI and UX Requirements

Match the application's existing design system exactly. Reuse its components,
spacing, and typography. This feature should look like it was always there.

## Responsive Requirements

Works on mobile, tablet, and desktop. Touch targets are large enough to hit on a
phone, and nothing overflows horizontally at 320px.

## Accessibility Requirements

- Fully keyboard navigable.
- Correct semantic elements and ARIA roles.
- Visible focus states.
- Meets WCAG AA contrast.
- Dynamic changes are announced to screen readers.
- Respects prefers-reduced-motion.

## Edge Cases

- An unusually large request must be priced before it is sent, not discovered afterwards. Estimate from the input size and refuse anything that alone would blow the remaining budget, telling the user what to trim.
- Streaming and tool-using runs do not have a known cost at dispatch. Reserve a conservative estimate up front, then reconcile against the real usage the provider reports when the run closes, releasing the difference.
- Several runs starting at the same moment must not each see the same remaining balance and all be admitted. Reserve atomically so concurrent requests cannot oversell the same headroom.
- Soft and hard limits must behave differently and visibly: a soft limit warns the owner and keeps working, a hard limit refuses new runs while letting in-flight ones finish and settle.
- The pricing figures used to compute spend change over time. Keep them versioned with effective dates, record which version priced each run, and never retroactively reprice settled history.
- If the provider never returns usage figures for a run, the reservation must still be settled on a timeout rather than pinning that amount forever.
- A budget reset at a period boundary must not release money for runs that are still in flight across it.
- When the model provider is unavailable, no budget is consumed. A failed dispatch must release its reservation rather than counting as spend.

## Testing

Exercise the feature end to end in the running application. Cover every edge case
above, then run the existing test suite and confirm nothing regressed.

## Acceptance Criteria

- [ ] Every model call in the app passes through a single accounting point that records spend against a scope.
- [ ] Large requests are cost-estimated before dispatch and refused with an explanation when they exceed the remaining budget.
- [ ] Reservations are made atomically and reconciled to actual usage, so concurrent runs cannot overspend a budget.
- [ ] Soft limits warn and continue while hard limits stop new work, and both are adjustable by an operator without a deploy.
- [ ] Pricing tables are versioned with effective dates and each run records the version that priced it.
- [ ] Failed, timed-out, and provider-unavailable runs release their reservations instead of counting as spend.
- [ ] The feature matches the existing design system.
- [ ] No existing functionality is broken.

## Adaptation Rules

- Match the existing design system. Do not introduce a new colour palette,
  spacing scale, or component library.
- Reuse existing components and utilities wherever they fit.
- Follow the naming, file layout, and code style already present.
- Do not upgrade, replace, or remove existing dependencies to make this
  feature fit. Adapt the feature to the app, not the app to the feature.
- Do not break existing functionality. If a change is genuinely required in
  existing code, make the smallest one that works and say so.
- If something in these instructions conflicts with how the application is
  built, follow the application and explain the deviation.

## Final Verification

Before you report the work as done:

1. Re-read the acceptance criteria above and check each one against what you
   actually built.
2. Run the application and exercise the feature end to end.
3. Run the existing test suite and confirm you have broken nothing.
4. Check the feature on mobile, tablet, and desktop widths.
5. Check keyboard navigation and focus handling.
6. Summarize what changed: files added, files modified, and anything you
   deliberately did differently because of how this application is built.

If any acceptance criterion is unmet, fix it before reporting completion.
