# AI Entity Extraction

## Objective

Pull names, dates, places, and amounts out of free text with the source spans kept.

Extraction of typed entities from unstructured text, each carrying its original wording, a normalized value, and its position in the source.

## Before You Begin

This feature is being added to an application that already exists and already
works. Do not scaffold a new project, and do not assume a blank slate.

Inspect the codebase first and establish:

- The existing application structure and where code of this kind already lives.
- The framework and version in use.
- The existing design system — colours, spacing, typography, and component conventions.
- Existing UI components you can reuse instead of writing new ones.
- The existing database structure, if this feature needs to persist anything.
- The existing authentication and authorization system, if this feature is user-scoped.
- Dependencies already installed, so you don't add a library that duplicates one.
- The existing test setup and conventions.

Only start writing code once you understand the above. If the application
already implements part of this feature, extend it rather than replacing it.

## Implementation Instructions

1. Fix the entity types up front — people, organizations, locations, products, dates, amounts, identifiers — and reject anything outside that set. An open-ended extractor produces a taxonomy nobody can query.
2. Require a source span for every entity: the offset and the exact original text. Without it a user cannot check the extraction, and a wrong value becomes indistinguishable from a right one.
3. Keep both forms of every value. Store the original wording as written and a normalized form beside it, so a display can show what the document said while a query can match across spellings.
4. Validate dates, currency amounts, phone numbers, and structured identifiers with the app's own deterministic parsers after extraction. The model is good at finding candidates and unreliable at formatting them.
5. This feature works within a single piece of text and returns spans. Assembling whole records against a schema belongs to AI Structured Data Extraction; have that feature call this one for candidate values rather than each maintaining its own extraction path.

## UI and UX Requirements

Match the application's existing design system exactly. Reuse its components,
spacing, and typography. This feature should look like it was always there.

## Responsive Requirements

Works on mobile, tablet, and desktop. Touch targets are large enough to hit on a
phone, and nothing overflows horizontally at 320px.

## Accessibility Requirements

- Fully keyboard navigable.
- Correct semantic elements and ARIA roles.
- Visible focus states.
- Meets WCAG AA contrast.
- Dynamic changes are announced to screen readers.
- Respects prefers-reduced-motion.

## Edge Cases

- Every entity must be traceable to a span in the source, and the interface should highlight it on request. An extraction that cannot be located in the original text must be dropped.
- Normalization must never overwrite the original. A company written three ways in one document needs one canonical value and three preserved surface forms, because the wording sometimes matters more than the match.
- The same entity appearing repeatedly should be returned once with all its occurrences, and genuinely ambiguous references should be returned as separate candidates rather than merged on a guess.
- Do not attach an extracted entity to an existing record automatically. Link only above a high confidence threshold, only where the current user may see the target record, and always with a visible way to undo.
- Run every date, number, and identifier through a real parser and discard whatever fails. A model-produced date that no parser accepts is not a date, and ambiguous day-month ordering must be resolved against the document's locale or left unresolved.
- Overlapping and nested spans occur naturally, such as a city inside an organization name. Define which wins and apply it consistently instead of returning both as peers.
- Long documents must be chunked with overlap so entities are not severed at a boundary, and offsets must be mapped back to the original document rather than to the chunk.
- When the provider fails or the response is truncated, save nothing partial. A half-extracted document presented as complete is worse than one marked unprocessed.

## Testing

Exercise the feature end to end in the running application. Cover every edge case
above, then run the existing test suite and confirm nothing regressed.

## Acceptance Criteria

- [ ] Every returned entity carries a type from the fixed set, a source span, an original form, and a normalized form.
- [ ] Dates, amounts, and identifiers are validated by deterministic parsers and invalid ones are discarded.
- [ ] Repeated mentions are grouped and ambiguous references remain separate candidates.
- [ ] Automatic linking to existing records respects confidence thresholds and the viewer's permissions, and is reversible.
- [ ] Chunked documents produce offsets correct against the original text.
- [ ] Failed or truncated runs leave the document marked unprocessed with no partial results stored.
- [ ] The feature matches the existing design system.
- [ ] No existing functionality is broken.

## Adaptation Rules

- Match the existing design system. Do not introduce a new colour palette,
  spacing scale, or component library.
- Reuse existing components and utilities wherever they fit.
- Follow the naming, file layout, and code style already present.
- Do not upgrade, replace, or remove existing dependencies to make this
  feature fit. Adapt the feature to the app, not the app to the feature.
- Do not break existing functionality. If a change is genuinely required in
  existing code, make the smallest one that works and say so.
- If something in these instructions conflicts with how the application is
  built, follow the application and explain the deviation.

## Final Verification

Before you report the work as done:

1. Re-read the acceptance criteria above and check each one against what you
   actually built.
2. Run the application and exercise the feature end to end.
3. Run the existing test suite and confirm you have broken nothing.
4. Check the feature on mobile, tablet, and desktop widths.
5. Check keyboard navigation and focus handling.
6. Summarize what changed: files added, files modified, and anything you
   deliberately did differently because of how this application is built.

If any acceptance criterion is unmet, fix it before reporting completion.
