# AI Audit Trail

## Objective

Keep a reviewable record of what the AI saw, produced, recommended, and actually did.

A per-run record of inputs, sources, configuration, output, approvals, and executed actions, restricted and retained by purpose.

## Before You Begin

This feature is being added to an application that already exists and already
works. Do not scaffold a new project, and do not assume a blank slate.

Inspect the codebase first and establish:

- The existing application structure and where code of this kind already lives.
- The framework and version in use.
- The existing design system — colours, spacing, typography, and component conventions.
- Existing UI components you can reuse instead of writing new ones.
- The existing database structure, if this feature needs to persist anything.
- The existing authentication and authorization system, if this feature is user-scoped.
- Dependencies already installed, so you don't add a library that duplicates one.
- The existing test setup and conventions.

Only start writing code once you understand the above. If the application
already implements part of this feature, extend it rather than replacing it.

## Implementation Instructions

1. Write one record per run covering the feature invoked, the configuration and prompt version in force, the sources retrieved, the tools called, the output produced, who approved it, and what was carried out.
2. Distinguish a draft from an executed action explicitly in the record. A trail that shows only what text was produced cannot answer the question that matters in a review, which is whether anything happened.
3. Redact at the moment of writing. Strip credentials, tokens, and payment details before the record is stored, and keep only the personal content genuinely needed to understand the decision.
4. Make records append-only, permission-restricted, and tamper-evident, and link out to the retrieval capture owned by Retrieval Debugger rather than duplicating chunk-level detail here. A trail the operator who caused the event can quietly edit proves nothing.
5. Set retention per purpose rather than one global window, since routine assistance, safety review, billing reconciliation, and regulated workflows justify different periods. Do not keep every prompt and response indefinitely because it was easier than deciding.

## UI and UX Requirements

Match the application's existing design system exactly. Reuse its components,
spacing, and typography. This feature should look like it was always there.

## Responsive Requirements

Works on mobile, tablet, and desktop. Touch targets are large enough to hit on a
phone, and nothing overflows horizontally at 320px.

## Accessibility Requirements

- Fully keyboard navigable.
- Correct semantic elements and ARIA roles.
- Visible focus states.
- Meets WCAG AA contrast.
- Dynamic changes are announced to screen readers.
- Respects prefers-reduced-motion.

## Edge Cases

- The record must capture the configuration and prompt version used, the sources supplied, the tools invoked, who approved the result, and the outcome. Any one of those missing makes an incident unreconstructable.
- Prompts and outputs carry secrets and personal data by accident. Redact before storage rather than at display time, because a redaction applied in the view still leaves the original in the database and in backups.
- A suggestion the user rejected, a draft never sent, and an action executed against live data must be visually and structurally distinct in the trail.
- Records must be append-only with modification detectable, and readable only by roles entitled to see the underlying content, since the trail aggregates material from across the app.
- Retention has to be set separately for ordinary usage, safety investigations, billing evidence, and regulated activity, because a single window is either too short for one of them or too long for the rest.
- A run that failed, timed out, was refused by the model, or was cancelled partway still needs a record. The absence of an entry must never be the only evidence that something went wrong.
- If writing the audit record fails, decide deliberately whether the action proceeds or is blocked, and apply the stricter choice to anything consequential.
- A deletion request from a user has to reach this trail as well. Define what is removed and what is retained under a legal obligation, and be able to state which is which.

## Testing

Exercise the feature end to end in the running application. Cover every edge case
above, then run the existing test suite and confirm nothing regressed.

## Acceptance Criteria

- [ ] Every AI run produces a record, including failed, refused, and cancelled runs.
- [ ] Each record identifies the configuration and prompt version, the sources retrieved, the tools called, and the approver.
- [ ] Drafts, rejected suggestions, and executed actions are distinguishable in the trail.
- [ ] Credentials and payment details are removed before the record is stored, not hidden at display time.
- [ ] Records are append-only, tamper-evident, and readable only by authorized roles.
- [ ] Retention periods are configured separately per purpose and enforced by deletion.
- [ ] Chunk-level retrieval detail is linked from Retrieval Debugger rather than duplicated.
- [ ] The feature matches the existing design system.
- [ ] No existing functionality is broken.

## Adaptation Rules

- Match the existing design system. Do not introduce a new colour palette,
  spacing scale, or component library.
- Reuse existing components and utilities wherever they fit.
- Follow the naming, file layout, and code style already present.
- Do not upgrade, replace, or remove existing dependencies to make this
  feature fit. Adapt the feature to the app, not the app to the feature.
- Do not break existing functionality. If a change is genuinely required in
  existing code, make the smallest one that works and say so.
- If something in these instructions conflicts with how the application is
  built, follow the application and explain the deviation.

## Final Verification

Before you report the work as done:

1. Re-read the acceptance criteria above and check each one against what you
   actually built.
2. Run the application and exercise the feature end to end.
3. Run the existing test suite and confirm you have broken nothing.
4. Check the feature on mobile, tablet, and desktop widths.
5. Check keyboard navigation and focus handling.
6. Summarize what changed: files added, files modified, and anything you
   deliberately did differently because of how this application is built.

If any acceptance criterion is unmet, fix it before reporting completion.
