# Data Anonymization

## Objective

Strip identity out of old records without wrecking your reporting.

An irreversible scrub of identifying fields across the schema that keeps rows, relationships, and aggregate counts intact.

## Before You Begin

This feature is being added to an application that already exists and already
works. Do not scaffold a new project, and do not assume a blank slate.

Inspect the codebase first and establish:

- The existing application structure and where code of this kind already lives.
- The framework and version in use.
- The existing design system — colours, spacing, typography, and component conventions.
- Existing UI components you can reuse instead of writing new ones.
- The existing database structure, if this feature needs to persist anything.
- The existing authentication and authorization system, if this feature is user-scoped.
- Dependencies already installed, so you don't add a library that duplicates one.
- The existing test setup and conventions.

Only start writing code once you understand the above. If the application
already implements part of this feature, extend it rather than replacing it.

## Implementation Instructions

1. Classify every column as identifying, quasi-identifying, or non-identifying. Free-text notes and file names carry identity too and are the ones teams forget.
2. Replace identifying values in place rather than deleting rows, so foreign keys stay valid and historical counts do not change.
3. Make the replacement irreversible: random or tokenized values with no stored mapping back to the original. If a lookup table exists, this is pseudonymization, not anonymization — call it that.
4. Extend the scrub beyond the primary tables to logs, search indexes, denormalized copies, cached aggregates, uploaded file contents and names, and analytics platforms.
5. Record that a given record was anonymized, when, and under which policy, without retaining the identity that was removed.
6. Do NOT use a deterministic transform such as hashing an email. The input space is small enough to enumerate, so the original is recoverable and the record is not anonymous.

## UI and UX Requirements

Match the application's existing design system exactly. Reuse its components,
spacing, and typography. This feature should look like it was always there.

## Responsive Requirements

Works on mobile, tablet, and desktop. Touch targets are large enough to hit on a
phone, and nothing overflows horizontally at 320px.

## Accessibility Requirements

- Fully keyboard navigable.
- Correct semantic elements and ARIA roles.
- Visible focus states.
- Meets WCAG AA contrast.
- Dynamic changes are announced to screen readers.
- Respects prefers-reduced-motion.

## Edge Cases

- Quasi-identifiers combine: a rare job title plus a postcode plus a signup date can identify one person even with the name removed.
- Uniqueness constraints will collide when several anonymized rows get the same placeholder — generate unique values or relax the constraint deliberately.
- Anonymizing a user who authored content must not break the content's display; render a stable neutral label rather than a blank byline.
- Aggregate reports must produce the same totals after anonymization as before — verify this, do not assume it.
- Anonymization is irreversible, so it needs the same dry-run and confirmation discipline as deletion.
- The scrub job must be resumable and must not leave a record half-anonymized across tables.
- Backups and prior exports still contain the original identity — state that explicitly rather than claiming the data is gone everywhere.

## Testing

Exercise the feature end to end in the running application. Cover every edge case
above, then run the existing test suite and confirm nothing regressed.

## Acceptance Criteria

- [ ] Every column is classified and every identifying column has a defined treatment.
- [ ] Records are scrubbed in place; row counts and aggregate totals are unchanged.
- [ ] Replacements are non-deterministic and no reversal mapping is retained.
- [ ] Logs, search indexes, denormalized fields, files, and third-party platforms are covered.
- [ ] An audit record proves anonymization occurred without storing the removed identity.
- [ ] Content authored by an anonymized user still renders correctly.
- [ ] The job is resumable and never leaves a record partially anonymized.
- [ ] The feature matches the existing design system.
- [ ] No existing functionality is broken.

## Adaptation Rules

- Match the existing design system. Do not introduce a new colour palette,
  spacing scale, or component library.
- Reuse existing components and utilities wherever they fit.
- Follow the naming, file layout, and code style already present.
- Do not upgrade, replace, or remove existing dependencies to make this
  feature fit. Adapt the feature to the app, not the app to the feature.
- Do not break existing functionality. If a change is genuinely required in
  existing code, make the smallest one that works and say so.
- If something in these instructions conflicts with how the application is
  built, follow the application and explain the deviation.

## Final Verification

Before you report the work as done:

1. Re-read the acceptance criteria above and check each one against what you
   actually built.
2. Run the application and exercise the feature end to end.
3. Run the existing test suite and confirm you have broken nothing.
4. Check the feature on mobile, tablet, and desktop widths.
5. Check keyboard navigation and focus handling.
6. Summarize what changed: files added, files modified, and anything you
   deliberately did differently because of how this application is built.

If any acceptance criterion is unmet, fix it before reporting completion.
