# Retrieval Debugger

## Objective

Show exactly which sources, chunks, and scores produced a given AI answer.

A per-answer inspector showing the query as issued, the filters applied, the candidate chunks with their scores, and what reached the model.

## Before You Begin

This feature is being added to an application that already exists and already
works. Do not scaffold a new project, and do not assume a blank slate.

Inspect the codebase first and establish:

- The existing application structure and where code of this kind already lives.
- The framework and version in use.
- The existing design system — colours, spacing, typography, and component conventions.
- Existing UI components you can reuse instead of writing new ones.
- The existing database structure, if this feature needs to persist anything.
- The existing authentication and authorization system, if this feature is user-scoped.
- Dependencies already installed, so you don't add a library that duplicates one.
- The existing test setup and conventions.

Only start writing code once you understand the above. If the application
already implements part of this feature, extend it rather than replacing it.

## Implementation Instructions

1. Capture for each answer the query as it was issued, the filters applied, the candidates returned with their scores, and which of those actually made it into the request after the context ceiling was applied.
2. Show results after permission filtering, with a count of how many candidates were excluded and why. Displaying the pre-filter set turns the debugger into a way to read content the viewer cannot open.
3. Present each scoring stage separately — keyword, semantic, and any reranking — because a chunk that ends up first overall may have been rescued by one stage after being buried by another, and a single blended number hides that.
4. Support replaying a captured query against the recorded index version, and label the replay clearly when the index has moved on since, so a comparison is never mistaken for the original run.
5. Do not attach debug metadata to normal end-user responses. Keep the whole surface behind operator permission, and leave index-wide counts and staleness to Vector Index Health rather than repeating them per answer.

## UI and UX Requirements

Match the application's existing design system exactly. Reuse its components,
spacing, and typography. This feature should look like it was always there.

## Responsive Requirements

Works on mobile, tablet, and desktop. Touch targets are large enough to hit on a
phone, and nothing overflows horizontally at 320px.

## Accessibility Requirements

- Fully keyboard navigable.
- Correct semantic elements and ARIA roles.
- Visible focus states.
- Meets WCAG AA contrast.
- Dynamic changes are announced to screen readers.
- Respects prefers-reduced-motion.

## Edge Cases

- Retrieval must be shown as the user experienced it, meaning after tenant and record filtering. A debugger that reports the unfiltered candidate set is reporting something that never influenced the answer.
- Chunk text the operator viewing the debugger has no right to read must be redacted to metadata — source, score, length — rather than displayed in full.
- Scores from different stages are not on the same scale and must not be plotted as though they were. Label each and show the ranking each stage produced.
- Replay needs the query, the filters, and the index version pinned together. Replaying against a changed index produces a different answer for reasons unrelated to the change being investigated.
- Debug records must never leak into ordinary responses, logs shown to customers, or error messages, since they contain retrieved content verbatim.
- Debug records hold customer content and therefore need their own retention window and deletion, shorter than the records they describe.
- Some answers are produced with no retrieval at all, and some with retrieval that returned nothing. Both need a clear representation rather than an empty table that looks like a capture failure.
- Replaying a query calls the model again and costs money, so it must be an explicit action with a visible warning, never something that happens on page load.

## Testing

Exercise the feature end to end in the running application. Cover every edge case
above, then run the existing test suite and confirm nothing regressed.

## Acceptance Criteria

- [ ] Every answer has a stored capture of its query, filters, candidates, scores, and final context.
- [ ] Results are displayed after permission filtering, with excluded candidates counted rather than shown.
- [ ] Content the viewing operator cannot access is redacted to metadata.
- [ ] Keyword, semantic, and rerank scores are shown separately and labelled.
- [ ] A query can be replayed against its recorded index version, with drift clearly marked.
- [ ] No debug metadata appears in end-user responses.
- [ ] Debug captures have their own retention period and are deleted when it expires.
- [ ] The feature matches the existing design system.
- [ ] No existing functionality is broken.

## Adaptation Rules

- Match the existing design system. Do not introduce a new colour palette,
  spacing scale, or component library.
- Reuse existing components and utilities wherever they fit.
- Follow the naming, file layout, and code style already present.
- Do not upgrade, replace, or remove existing dependencies to make this
  feature fit. Adapt the feature to the app, not the app to the feature.
- Do not break existing functionality. If a change is genuinely required in
  existing code, make the smallest one that works and say so.
- If something in these instructions conflicts with how the application is
  built, follow the application and explain the deviation.

## Final Verification

Before you report the work as done:

1. Re-read the acceptance criteria above and check each one against what you
   actually built.
2. Run the application and exercise the feature end to end.
3. Run the existing test suite and confirm you have broken nothing.
4. Check the feature on mobile, tablet, and desktop widths.
5. Check keyboard navigation and focus handling.
6. Summarize what changed: files added, files modified, and anything you
   deliberately did differently because of how this application is built.

If any acceptance criterion is unmet, fix it before reporting completion.
