# Scheduled Job Monitor

## Objective

Find out a nightly task stopped running before your users do.

A record of every recurring schedule, the runs it was expected to make, and alerts when a run is missed or late.

## Before You Begin

This feature is being added to an application that already exists and already
works. Do not scaffold a new project, and do not assume a blank slate.

Inspect the codebase first and establish:

- The existing application structure and where code of this kind already lives.
- The framework and version in use.
- The existing design system — colours, spacing, typography, and component conventions.
- Existing UI components you can reuse instead of writing new ones.
- The existing database structure, if this feature needs to persist anything.
- The existing authentication and authorization system, if this feature is user-scoped.
- Dependencies already installed, so you don't add a library that duplicates one.
- The existing test setup and conventions.

Only start writing code once you understand the above. If the application
already implements part of this feature, extend it rather than replacing it.

## Implementation Instructions

1. Register every recurring task with its schedule expression, its time zone, and an expected maximum duration. A schedule that is not registered cannot be monitored.
2. Record the start and the completion of each run separately. A task that started and never finished is a different failure from one that never started, and they need different alerts.
3. Compute the expected run times forward from the schedule and compare them against actual starts. Alert on a missed run, and alert separately when a run starts on time but overruns its expected duration.
4. Deduplicate alerts per schedule. One ongoing outage is one incident — send a notification when it opens, optionally a reminder, and one when it recovers. Do not fire an alert per missed run.
5. Support marking a schedule as intentionally paused, with who paused it and why, and suppress its alerts while paused.
6. Do NOT alert on the first late run of a schedule that has never successfully run. A newly deployed task with a wrong expression should surface as unconfigured, not as an outage.
7. Live queue state, payload inspection, and manual retry are owned by Background Job Dashboard; link across to it rather than duplicating queue controls here.

## UI and UX Requirements

Match the application's existing design system exactly. Reuse its components,
spacing, and typography. This feature should look like it was always there.

## Responsive Requirements

Works on mobile, tablet, and desktop. Touch targets are large enough to hit on a
phone, and nothing overflows horizontally at 320px.

## Accessibility Requirements

- Fully keyboard navigable.
- Correct semantic elements and ARIA roles.
- Visible focus states.
- Meets WCAG AA contrast.
- Dynamic changes are announced to screen readers.
- Respects prefers-reduced-motion.

## Edge Cases

- Daylight-saving transitions delete one local hour and repeat another. A schedule set for that hour must not silently skip or double-fire, and the monitor must not report the skip as a miss.
- Store and compare in UTC, but display in the schedule's own time zone. An operator debugging a 02:00 job needs to see 02:00.
- A schedule changed mid-week invalidates the historical expectation. Version the schedule and evaluate each past run against the expression in force at the time.
- The monitor itself can fail. If the checking process stops, nothing is reported as missed — treat a stale monitor as an alertable condition of its own.
- A run that overlaps the next scheduled run needs a policy: skip, queue, or run concurrently. Show which one applies.
- Failure reasons shown on the dashboard must be redacted the same way job payloads are — a stack trace from a nightly export routinely quotes customer data.
- A schedule deleted from the code must not alert forever. Mark unregistered schedules as retired rather than missing.

## Testing

Exercise the feature end to end in the running application. Cover every edge case
above, then run the existing test suite and confirm nothing regressed.

## Acceptance Criteria

- [ ] Every recurring task is registered with its schedule, time zone, and expected duration.
- [ ] Starts and completions are recorded separately, and overruns alert distinctly from misses.
- [ ] Each schedule shows last success, next expected run, last duration, and last failure reason.
- [ ] One ongoing incident produces one alert plus a recovery notification.
- [ ] Paused schedules suppress alerts and record who paused them.
- [ ] DST transitions produce neither a false miss nor a double run.
- [ ] A stopped monitoring process is itself detectable.
- [ ] The feature matches the existing design system.
- [ ] No existing functionality is broken.

## Adaptation Rules

- Match the existing design system. Do not introduce a new colour palette,
  spacing scale, or component library.
- Reuse existing components and utilities wherever they fit.
- Follow the naming, file layout, and code style already present.
- Do not upgrade, replace, or remove existing dependencies to make this
  feature fit. Adapt the feature to the app, not the app to the feature.
- Do not break existing functionality. If a change is genuinely required in
  existing code, make the smallest one that works and say so.
- If something in these instructions conflicts with how the application is
  built, follow the application and explain the deviation.

## Final Verification

Before you report the work as done:

1. Re-read the acceptance criteria above and check each one against what you
   actually built.
2. Run the application and exercise the feature end to end.
3. Run the existing test suite and confirm you have broken nothing.
4. Check the feature on mobile, tablet, and desktop widths.
5. Check keyboard navigation and focus handling.
6. Summarize what changed: files added, files modified, and anything you
   deliberately did differently because of how this application is built.

If any acceptance criterion is unmet, fix it before reporting completion.
