# AI Sentiment Analysis

## Objective

Estimate the tone of feedback and messages without pretending the model is certain.

A sentiment label with a confidence value attached to feedback, reviews, tickets, and messages the app already stores.

## Before You Begin

This feature is being added to an application that already exists and already
works. Do not scaffold a new project, and do not assume a blank slate.

Inspect the codebase first and establish:

- The existing application structure and where code of this kind already lives.
- The framework and version in use.
- The existing design system — colours, spacing, typography, and component conventions.
- Existing UI components you can reuse instead of writing new ones.
- The existing database structure, if this feature needs to persist anything.
- The existing authentication and authorization system, if this feature is user-scoped.
- Dependencies already installed, so you don't add a library that duplicates one.
- The existing test setup and conventions.

Only start writing code once you understand the above. If the application
already implements part of this feature, extend it rather than replacing it.

## Implementation Instructions

1. Find the free-text surfaces where tone would change how someone triages work: reviews, support messages, survey responses, comments. Score those on write and on demand, not the entire corpus in one pass.
2. Store the label, the confidence, the text version it was derived from, and the time it was produced. A score with no provenance cannot be audited or recomputed when the prompt changes.
3. Allow four outcomes plus a fifth: positive, neutral, negative, mixed, and unknown. A message that praises the product and condemns the billing is genuinely mixed, and forcing it into one bucket loses the part someone needs to act on.
4. Run scoring through the app's existing background-job system with a per-workspace token ceiling and a maximum input length, truncating on a sentence boundary rather than mid-word. Record spend per run so the cost of the feature is visible before it becomes a surprise.
5. Do not let a sentiment score gate anything consequential on its own. It may sort, filter, and flag for a human, but it must never close a ticket, downgrade an account, or suppress a message without review.

## UI and UX Requirements

Match the application's existing design system exactly. Reuse its components,
spacing, and typography. This feature should look like it was always there.

## Responsive Requirements

Works on mobile, tablet, and desktop. Touch targets are large enough to hit on a
phone, and nothing overflows horizontally at 320px.

## Accessibility Requirements

- Fully keyboard navigable.
- Correct semantic elements and ARIA roles.
- Visible focus states.
- Meets WCAG AA contrast.
- Dynamic changes are announced to screen readers.
- Respects prefers-reduced-motion.

## Edge Cases

- Sentiment describes one piece of text at one moment, not the person who wrote it. Do not roll scores up into a standing label on a customer record or display anything that reads as a judgement of the author.
- A single message can carry conflicting tone across its parts. Support a mixed outcome, and where the surface allows it, indicate which passages pulled which way rather than averaging them into a bland neutral.
- Sarcasm, one-word replies, emoji-only messages, and heavy in-house jargon defeat tone detection. Return unknown with low confidence in these cases and show nothing rather than showing a confident wrong answer.
- Automated decisions built on a probabilistic label will be wrong at scale. Keep escalation, closure, refunds, and account actions behind a human who can see the underlying text.
- The same words carry different weight across languages, regions, and product domains, so a threshold tuned on English support tickets will misread everything else. Calibrate per language and per surface, and hold back the feature where it has not been calibrated.
- When the model is unavailable, times out, or refuses, leave the record unscored and show it as not yet analysed. Do not fall back to neutral, which is indistinguishable from a real result.
- Malformed or unexpected output must be discarded rather than coerced. One retry with a lower ceiling is reasonable; a second failure means the record stays unscored.
- Redact or exclude payment details, credentials, and identity documents before any text leaves the app, and give operators a way to see exactly which fields are sent.

## Testing

Exercise the feature end to end in the running application. Cover every edge case
above, then run the existing test suite and confirm nothing regressed.

## Acceptance Criteria

- [ ] Every scored record stores its label, confidence, source text version, and timestamp.
- [ ] Mixed and unknown are first-class outcomes and appear in the interface.
- [ ] Low-confidence results are withheld from display rather than shown as fact.
- [ ] No account, ticket, or moderation action is taken from a sentiment score without human review.
- [ ] Provider outages, timeouts, and refusals leave records visibly unscored instead of defaulting to neutral.
- [ ] Token and cost ceilings are enforced per workspace and spend is reportable.
- [ ] The feature matches the existing design system.
- [ ] No existing functionality is broken.

## Adaptation Rules

- Match the existing design system. Do not introduce a new colour palette,
  spacing scale, or component library.
- Reuse existing components and utilities wherever they fit.
- Follow the naming, file layout, and code style already present.
- Do not upgrade, replace, or remove existing dependencies to make this
  feature fit. Adapt the feature to the app, not the app to the feature.
- Do not break existing functionality. If a change is genuinely required in
  existing code, make the smallest one that works and say so.
- If something in these instructions conflicts with how the application is
  built, follow the application and explain the deviation.

## Final Verification

Before you report the work as done:

1. Re-read the acceptance criteria above and check each one against what you
   actually built.
2. Run the application and exercise the feature end to end.
3. Run the existing test suite and confirm you have broken nothing.
4. Check the feature on mobile, tablet, and desktop widths.
5. Check keyboard navigation and focus handling.
6. Summarize what changed: files added, files modified, and anything you
   deliberately did differently because of how this application is built.

If any acceptance criterion is unmet, fix it before reporting completion.
