# Document Question Answering

## Objective

Let users ask questions about a document and get answers drawn only from that document.

An extraction and retrieval pipeline over uploaded files, with answers that cite the passages they came from.

## Before You Begin

This feature is being added to an application that already exists and already
works. Do not scaffold a new project, and do not assume a blank slate.

Inspect the codebase first and establish:

- The existing application structure and where code of this kind already lives.
- The framework and version in use.
- The existing design system — colours, spacing, typography, and component conventions.
- Existing UI components you can reuse instead of writing new ones.
- The existing database structure, if this feature needs to persist anything.
- The existing authentication and authorization system, if this feature is user-scoped.
- Dependencies already installed, so you don't add a library that duplicates one.
- The existing test setup and conventions.

Only start writing code once you understand the above. If the application
already implements part of this feature, extend it rather than replacing it.

## Implementation Instructions

1. Extract text from each upload with the app's existing background-job system, and record per-page or per-section extraction status so unreadable parts are known rather than silently missing.
2. Retrieve candidate passages first, then answer only from those passages. Instruct the model to say the document does not cover a question rather than filling the gap from general knowledge.
3. Attach a citation to every claim, pointing at the specific passage and its location in the file, and make the citation open that location in the document viewer.
4. Isolate extracted text and any derived index by tenant and by the document's own permissions, and re-check permission at query time rather than relying on the index being correctly partitioned.
5. When a document is replaced or edited, mark its index stale immediately and block answers until re-extraction finishes. Do not serve answers from the previous version's passages under the new file's name.

## UI and UX Requirements

Match the application's existing design system exactly. Reuse its components,
spacing, and typography. This feature should look like it was always there.

## Responsive Requirements

Works on mobile, tablet, and desktop. Touch targets are large enough to hit on a
phone, and nothing overflows horizontally at 320px.

## Accessibility Requirements

- Fully keyboard navigable.
- Correct semantic elements and ARIA roles.
- Visible focus states.
- Meets WCAG AA contrast.
- Dynamic changes are announced to screen readers.
- Respects prefers-reduced-motion.

## Edge Cases

- Scanned pages, images without text, password-protected files, and unsupported formats will appear. Report exactly which pages or sections could not be read and keep answering from the ones that could.
- An answer with no citation is not trustworthy. If no retrieved passage supports the claim, return the no-answer state rather than an uncited paragraph.
- The document not containing the answer is a correct outcome, not a failure. Say so plainly and offer the closest passages found, without inventing a synthesis.
- Extracted text and any index derived from it inherit the document's access rules. A user who loses access to the file must immediately stop getting answers built from it.
- Replacing a document must invalidate cached answers and prior citations. A citation pointing at a page number that no longer exists must be shown as expired, not followed blindly.
- Very large documents will exceed both the extraction budget and the context window. Cap the file size accepted and tell the user the limit before they wait for an upload to fail.
- If the model is unavailable, keep the document viewer and its full-text search working so the file is still usable.
- Knowledge Base Chat answers across the app's published sources. This feature answers within a single user-supplied document; do not mix the two corpora in one answer.

## Testing

Exercise the feature end to end in the running application. Cover every edge case
above, then run the existing test suite and confirm nothing regressed.

## Acceptance Criteria

- [ ] Extraction runs in the background and records which pages or sections were unreadable.
- [ ] Every answer carries citations that resolve to a location in the source document.
- [ ] Questions the document does not cover return an explicit no-answer response rather than a generated one.
- [ ] Extracted text and derived indexes are unreachable across tenants, and permission is verified at query time.
- [ ] Replacing a document invalidates its index and its cached answers before any new question is answered.
- [ ] Document viewing and keyword search continue to work when the model is unavailable.
- [ ] The feature matches the existing design system.
- [ ] No existing functionality is broken.

## Adaptation Rules

- Match the existing design system. Do not introduce a new colour palette,
  spacing scale, or component library.
- Reuse existing components and utilities wherever they fit.
- Follow the naming, file layout, and code style already present.
- Do not upgrade, replace, or remove existing dependencies to make this
  feature fit. Adapt the feature to the app, not the app to the feature.
- Do not break existing functionality. If a change is genuinely required in
  existing code, make the smallest one that works and say so.
- If something in these instructions conflicts with how the application is
  built, follow the application and explain the deviation.

## Final Verification

Before you report the work as done:

1. Re-read the acceptance criteria above and check each one against what you
   actually built.
2. Run the application and exercise the feature end to end.
3. Run the existing test suite and confirm you have broken nothing.
4. Check the feature on mobile, tablet, and desktop widths.
5. Check keyboard navigation and focus handling.
6. Summarize what changed: files added, files modified, and anything you
   deliberately did differently because of how this application is built.

If any acceptance criterion is unmet, fix it before reporting completion.
