# Experiment Results Guardrails

## Objective

Say whether a test can be trusted yet, before anyone declares a winner.

A results view that reports sample sufficiency, exposure quality, and effect size alongside the numbers — and refuses to crown a winner early.

## Before You Begin

This feature is being added to an application that already exists and already
works. Do not scaffold a new project, and do not assume a blank slate.

Inspect the codebase first and establish:

- The existing application structure and where code of this kind already lives.
- The framework and version in use.
- The existing design system — colours, spacing, typography, and component conventions.
- Existing UI components you can reuse instead of writing new ones.
- The existing database structure, if this feature needs to persist anything.
- The existing authentication and authorization system, if this feature is user-scoped.
- Dependencies already installed, so you don't add a library that duplicates one.
- The existing test setup and conventions.

Only start writing code once you understand the above. If the application
already implements part of this feature, extend it rather than replacing it.

## Implementation Instructions

1. Compute results from exposure events — users who actually saw the variant — not from assignment counts. Assigned-but-never-exposed users dilute every rate.
2. Fix the primary metric and the stopping rule when the experiment is created, and version them. A metric changed after the data arrives is not a result.
3. Show a confidence interval and the absolute effect next to every relative lift. A 40% improvement on a 0.1% baseline is noise dressed as a win.
4. Compute the sample size needed for the stated minimum detectable effect, and show progress towards it prominently.
5. Do NOT display a winner, a green badge, or a significance verdict before the stopping rule is met. Show the guardrail instead and say why.

## UI and UX Requirements

Match the application's existing design system exactly. Reuse its components,
spacing, and typography. This feature should look like it was always there.

## Responsive Requirements

Works on mobile, tablet, and desktop. Touch targets are large enough to hit on a
phone, and nothing overflows horizontally at 320px.

## Accessibility Requirements

- Fully keyboard navigable.
- Correct semantic elements and ARIA roles.
- Visible focus states.
- Meets WCAG AA contrast.
- Dynamic changes are announced to screen readers.
- Respects prefers-reduced-motion.

## Edge Cases

- Sample ratio mismatch — buckets arriving at meaningfully different sizes — means assignment is broken. Detect it and invalidate the readout rather than interpreting the numbers.
- Repeated checking inflates false positives. Either record how many times results were viewed and warn, or use a method that tolerates peeking.
- The current day and the first hours of an experiment are partial. Exclude or mark them rather than letting them swing the totals.
- Weekday and weekend traffic behave differently. Warn when an experiment has not covered whole weekly cycles.
- Testing many secondary metrics will find something significant by chance. Label secondary metrics as exploratory and adjust or say you have not.
- A guardrail metric moving the wrong way — errors, latency, refunds — must be surfaced even when the primary metric wins.

## Testing

Exercise the feature end to end in the running application. Cover every edge case
above, then run the existing test suite and confirm nothing regressed.

## Acceptance Criteria

- [ ] Rates are computed from exposure events, not assignment records.
- [ ] The primary metric and stopping rule are fixed at creation and versioned on change.
- [ ] Required sample size and progress towards it are shown on the results view.
- [ ] Absolute effect and a confidence interval accompany every relative lift.
- [ ] Sample ratio mismatch is detected and blocks the readout.
- [ ] No winner is declared before the stopping rule is satisfied.
- [ ] Guardrail metrics are shown regardless of the primary result.
- [ ] The feature matches the existing design system.
- [ ] No existing functionality is broken.

## Adaptation Rules

- Match the existing design system. Do not introduce a new colour palette,
  spacing scale, or component library.
- Reuse existing components and utilities wherever they fit.
- Follow the naming, file layout, and code style already present.
- Do not upgrade, replace, or remove existing dependencies to make this
  feature fit. Adapt the feature to the app, not the app to the feature.
- Do not break existing functionality. If a change is genuinely required in
  existing code, make the smallest one that works and say so.
- If something in these instructions conflicts with how the application is
  built, follow the application and explain the deviation.

## Final Verification

Before you report the work as done:

1. Re-read the acceptance criteria above and check each one against what you
   actually built.
2. Run the application and exercise the feature end to end.
3. Run the existing test suite and confirm you have broken nothing.
4. Check the feature on mobile, tablet, and desktop widths.
5. Check keyboard navigation and focus handling.
6. Summarize what changed: files added, files modified, and anything you
   deliberately did differently because of how this application is built.

If any acceptance criterion is unmet, fix it before reporting completion.
