AddThisFeature

Scheduled Job Monitor

Find out a nightly task stopped running before your users do.

moderate Admin & Operations

What it adds

A record of every recurring schedule, the runs it was expected to make, and alerts when a run is missed or late.

What your agent is told to do

7
  1. 1

    Register every recurring task with its schedule expression, its time zone, and an expected maximum duration. A schedule that is not registered cannot be monitored.

  2. 2

    Record the start and the completion of each run separately. A task that started and never finished is a different failure from one that never started, and they need different alerts.

  3. 3

    Compute the expected run times forward from the schedule and compare them against actual starts. Alert on a missed run, and alert separately when a run starts on time but overruns its expected duration.

  4. 4

    Deduplicate alerts per schedule. One ongoing outage is one incident — send a notification when it opens, optionally a reminder, and one when it recovers. Do not fire an alert per missed run.

  5. 5

    Support marking a schedule as intentionally paused, with who paused it and why, and suppress its alerts while paused.

  6. 6

    Do NOT alert on the first late run of a schedule that has never successfully run. A newly deployed task with a wrong expression should surface as unconfigured, not as an outage.

  7. 7

    Live queue state, payload inspection, and manual retry are owned by Background Job Dashboard; link across to it rather than duplicating queue controls here.

Edge cases it handles

7
  • Daylight-saving transitions delete one local hour and repeat another. A schedule set for that hour must not silently skip or double-fire, and the monitor must not report the skip as a miss.
  • Store and compare in UTC, but display in the schedule's own time zone. An operator debugging a 02:00 job needs to see 02:00.
  • A schedule changed mid-week invalidates the historical expectation. Version the schedule and evaluate each past run against the expression in force at the time.
  • The monitor itself can fail. If the checking process stops, nothing is reported as missed — treat a stale monitor as an alertable condition of its own.
  • A run that overlaps the next scheduled run needs a policy: skip, queue, or run concurrently. Show which one applies.
  • Failure reasons shown on the dashboard must be redacted the same way job payloads are — a stack trace from a nightly export routinely quotes customer data.
  • A schedule deleted from the code must not alert forever. Mark unregistered schedules as retired rather than missing.

Definition of done

9
  • Every recurring task is registered with its schedule, time zone, and expected duration.
  • Starts and completions are recorded separately, and overruns alert distinctly from misses.
  • Each schedule shows last success, next expected run, last duration, and last failure reason.
  • One ongoing incident produces one alert plus a recovery notification.
  • Paused schedules suppress alerts and record who paused them.
  • DST transitions produce neither a false miss nor a double run.
  • A stopped monitoring process is itself detectable.
  • The feature matches the existing design system.
  • No existing functionality is broken.

Related features

How it works

  1. 1

    Copy the link

    Grab the Markdown instruction URL for this feature.

  2. 2

    Give it to your AI

    Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.

  3. 3

    It inspects, then implements

    Your agent reads your existing app first, then adds the feature to fit it.

Works with your stack

These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.

Need it tighter than that? Customize the feature and tell it exactly what you're running.