# AI Tool Calling Framework

## Objective

Give the model a governed catalogue of app actions instead of ad hoc, unchecked calls.

A registry of app capabilities exposed to the model, with per-user scoping, strict argument validation, execution limits, and an audit trail.

## Before You Begin

This feature is being added to an application that already exists and already
works. Do not scaffold a new project, and do not assume a blank slate.

Inspect the codebase first and establish:

- The existing application structure and where code of this kind already lives.
- The framework and version in use.
- The existing design system — colours, spacing, typography, and component conventions.
- Existing UI components you can reuse instead of writing new ones.
- The existing database structure, if this feature needs to persist anything.
- The existing authentication and authorization system, if this feature is user-scoped.
- Dependencies already installed, so you don't add a library that duplicates one.
- The existing test setup and conventions.

Only start writing code once you understand the above. If the application
already implements part of this feature, extend it rather than replacing it.

## Implementation Instructions

1. Define each tool once in a registry with its purpose, its argument schema, the permission it requires, and whether it reads or writes. Anything not in the registry is not callable.
2. Build the tool list per request from the current user, workspace, and context. Do not publish the full catalogue and rely on the model to avoid the tools it should not use.
3. Validate every argument against the schema and reject the call on any deviation. Then apply the same authorization checks the ordinary interface applies, in the app's own code, because a tool call is an untrusted request that happens to arrive from a model.
4. Enforce ceilings on the number of calls per turn, recursion depth, wall-clock time, and token or cost spend. Stop cleanly at the ceiling with an explanation rather than looping until something times out.
5. Do not let a write tool execute directly when a user-visible plan is warranted. Approval, plan display, and partial-failure reporting are owned by AI Agent Actions with Approval; this framework owns the catalogue, validation, and limits it runs on.

## UI and UX Requirements

Match the application's existing design system exactly. Reuse its components,
spacing, and typography. This feature should look like it was always there.

## Responsive Requirements

Works on mobile, tablet, and desktop. Touch targets are large enough to hit on a
phone, and nothing overflows horizontally at 320px.

## Accessibility Requirements

- Fully keyboard navigable.
- Correct semantic elements and ARIA roles.
- Visible focus states.
- Meets WCAG AA contrast.
- Dynamic changes are announced to screen readers.
- Respects prefers-reduced-motion.

## Edge Cases

- The tool list is scoped per request. A user who loses access mid-conversation must not be able to call a tool that was available earlier in the same thread.
- Reject any call whose arguments fail the schema, including extra fields, wrong types, or out-of-range values, and return the rejection to the model as a correctable error rather than coercing the value into something plausible.
- Tool results are untrusted input. Text pulled from a record, a document, or a third party may contain instructions aimed at the model, so it must be clearly delimited as data and must never be able to grant permissions or expand the tool list.
- Ceilings on call count, recursion, execution time, and spend must all be enforced, and the loop must terminate visibly at whichever is hit first. A model that calls the same read tool repeatedly is the normal failure, not the rare one.
- Record every call, its arguments, its outcome, and its duration, with secrets, tokens, and personal data redacted before anything is written. The log exists to explain what happened, not to reproduce the payload.
- A tool that fails, times out, or is unavailable must return a structured failure the model can reason about, rather than an exception that ends the conversation.
- Write tools need idempotency keys, because a model that does not see a result will call again.
- Registry changes are a compatibility surface. Removing or renaming a tool must not break conversations already in flight.

## Testing

Exercise the feature end to end in the running application. Cover every edge case
above, then run the existing test suite and confirm nothing regressed.

## Acceptance Criteria

- [ ] Only registered tools are callable, and the list offered is built per request from the current user and context.
- [ ] Every argument is validated against a strict schema, and failing calls are rejected rather than coerced.
- [ ] Authorization is enforced in application code on every call, independently of the tool list.
- [ ] Call count, recursion depth, execution time, and spend ceilings are enforced and terminate the loop cleanly.
- [ ] Tool output is treated as untrusted data and cannot expand the model's permissions or tool list.
- [ ] Every call and outcome is recorded with secrets and personal data redacted.
- [ ] Tool failures return a structured result rather than ending the conversation.
- [ ] The feature matches the existing design system.
- [ ] No existing functionality is broken.

## Adaptation Rules

- Match the existing design system. Do not introduce a new colour palette,
  spacing scale, or component library.
- Reuse existing components and utilities wherever they fit.
- Follow the naming, file layout, and code style already present.
- Do not upgrade, replace, or remove existing dependencies to make this
  feature fit. Adapt the feature to the app, not the app to the feature.
- Do not break existing functionality. If a change is genuinely required in
  existing code, make the smallest one that works and say so.
- If something in these instructions conflicts with how the application is
  built, follow the application and explain the deviation.

## Final Verification

Before you report the work as done:

1. Re-read the acceptance criteria above and check each one against what you
   actually built.
2. Run the application and exercise the feature end to end.
3. Run the existing test suite and confirm you have broken nothing.
4. Check the feature on mobile, tablet, and desktop widths.
5. Check keyboard navigation and focus handling.
6. Summarize what changed: files added, files modified, and anything you
   deliberately did differently because of how this application is built.

If any acceptance criterion is unmet, fix it before reporting completion.
