> ## Documentation Index
> Fetch the complete documentation index at: https://docs.heylua.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Watch schedule health

> Declare health on a scheduled workflow or job so the platform tells its owner when it goes silent, keeps failing or is turned off, and closes the alert when it recovers

After this guide, a scheduled [workflow](/concepts/workflows) or [job](/concepts/jobs) says how it is supposed to behave, and the platform tells its owner when it stops: when no scheduled run has succeeded for too long, when runs keep failing, or when the platform has turned the schedule off. You do not need a heartbeat service, a cron monitor or an alert of your own.

*Compiling `health` needs the next lua-cli release. A lua-cli that does not know the field drops it without an error, and the schedule is pushed unwatched.*

**Before you begin**

* A workflow with a `schedule` ([Schedule a workflow and set goals](/build/workflows/goals-and-schedules)) or a `LuaJob` with a recurring `schedule` ([Schedule a recurring job](/build/schedule-a-job)).

<Steps>
  <Step title="Declare health">
    `health` sits beside `schedule`, on `createWorkflow` or on `LuaJob`.

    ```ts src/workflows/heartbeat.ts theme={null}
    import { createWorkflow } from 'lua-cli';

    export const heartbeat = createWorkflow({
      name: 'heartbeat',
      schedule: { type: 'cron', expression: '*/10 * * * *', timezone: 'UTC' },
      scheduleInput: {},
      health: { maxSilenceMinutes: 30, alertAfterConsecutiveFailures: 3, notify: 'emailApp' },
    })
      .then(beat)
      .commit();
    ```

    ```ts src/jobs/NightlySyncJob.ts theme={null}
    import { LuaJob } from 'lua-cli';

    export default new LuaJob({
      name: 'nightly-sync',
      description: 'Sync the catalogue',
      schedule: { type: 'cron', expression: '0 2 * * *' },
      health: { maxSilenceMinutes: 26 * 60, alertAfterConsecutiveFailures: 2 },
      async execute() {
        return sync();
      },
    });
    ```

    | Field                           | Range                                  | Opens an incident when, while the schedule is active,                               |
    | ------------------------------- | -------------------------------------- | ----------------------------------------------------------------------------------- |
    | `maxSilenceMinutes`             | 5 to 43,200 (30 days)                  | no scheduled run has completed successfully for this many minutes                   |
    | `alertAfterConsecutiveFailures` | Workflows 1 to 3; jobs 1 to 50         | this many scheduled runs in a row have failed                                       |
    | `notify`                        | `app`, `email` or `emailApp` (default) | Not a condition: where the notice goes beyond the inbox card (push, email, or both) |

    Give at least one of the two thresholds. A schedule the platform **turns off** after repeated failures is always an incident while it stays off, which is why a workflow's failure threshold stops at 3: the platform turns a workflow schedule off after three failed runs in a row, so a fourth never happens.

    A cron cadence is not parsed, so leave headroom above the cron period plus the run time: `*/10` with a 5-minute run wants 30, not 15.
  </Step>

  <Step title="Compile and push">
    `lua compile` checks the declaration and `lua push` checks it again. Both refuse, with the code shown:

    | Code                                                                                                 | When                                                          |
    | ---------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- |
    | `health-empty`                                                                                       | Neither threshold is given                                    |
    | `health-requires-schedule`                                                                           | There is no `schedule`                                        |
    | `health-silence-on-once`                                                                             | Silence is declared on a one-time schedule                    |
    | `health-silence-shorter-than-interval`                                                               | The window is not longer than an `interval` schedule's period |
    | `health-failures-unreachable`                                                                        | A workflow's failure threshold is above 3                     |
    | `health-silence-invalid`, `health-failures-invalid`, `health-notify-invalid`, `health-unknown-field` | A value is out of range, or a key is not a health field       |

    The declaration takes effect when the version is live: for a workflow when its version is published (`lua push` and then activate), for a job when its version is deployed or promoted. Publishing a version without `health` stops the watch, and an alert that was open for it is closed.
  </Step>

  <Step title="Know what happens">
    The platform checks each schedule that declares `health` every few minutes.

    1. **An incident opens.** The owner receives one notice: the schedule's creator, or the organization's admins for a schedule nobody owns. It is an inbox card titled with the schedule's name, plus push or email as `notify` says. It is a card of its own, separate from the schedule's run and auto-disable cards.
    2. **Later checks stay quiet.** Nobody is pinged again while the same incident is open. If the condition changes, for example from failures to the platform turning the schedule off, the open incident is updated in place. A notice that could not be delivered is sent again on the next check and still counts as the same notice.
    3. **The incident resolves** when the schedule is healthy again: a scheduled run completes inside the window and the failure streak is broken. The same card closes ("Schedule healthy again") and a recovery notice goes out on the same channels. Pausing the schedule also closes it ("Schedule paused — health alert closed"). Resuming gives the schedule a full window before it can be called silent again.

    What counts:

    * **Workflows.** Only runs fired by the schedule. A manual `lua workflows start` is neither a success nor a failure. Failed, timed-out and abandoned runs, and fires that could not start, are failures; a cancelled run counts as neither.
    * **Jobs.** Every execution of the job, including `lua jobs trigger`. `failed`, `timeout`, `killed`, `abandoned` and `reaped` are failures; cancelled and still-running executions are skipped.
  </Step>

  <Step title="See it">
    ```bash theme={null}
    lua workflows schedules list
    lua workflows schedules list --json
    lua jobs view -i nightly-sync
    ```

    `lua workflows schedules list` has a **Health** column and `lua jobs view` a **Health:** line: `—` when nothing is declared, `ok`, `unwatched`, or `ALERT <condition> since <time>`. With `--json`, each schedule carries `health: { config, state, condition?, since?, lastSucceededAt?, checkedAt? }` and `lastSucceededAt`. `state` is `ok` when the schedule is watched and healthy, `alerting` while an incident is open, and `unwatched` while the schedule is paused or before its first check, up to about 5 minutes after publishing. `condition` is `silence`, `failures` or `auto_disabled`.
  </Step>

  <Step title="Route alerts to your own tools">
    Every open, change and resolve is also a line in the agent's logs: `lua logs --type workflow`, or `--type job` for a job. A `warn` line marks an incident and an `info` line its resolution. Its fields are `health_event`, `health_condition`, `job_id`, `consecutive_failures`, `last_succeeded_at`, `silent_for_minutes`, `max_silence_minutes` and `alert_after_consecutive_failures`.

    ```bash theme={null}
    lua logs --type workflow --limit 20
    ```

    A [log drain](/drains/overview) carries those lines like any other, so Datadog, Better Stack or your SIEM can alert on `health_event` `opened` for every agent in the organization with no per-agent setup.
  </Step>
</Steps>

## Options you may need

### Resume a schedule the platform turned off

After three failed runs in a row the platform turns a workflow schedule off. Reactivating the workflow does not turn it back on. Fix the cause, then resume the schedule by its job id from `lua workflows schedules list`:

```bash theme={null}
lua workflows schedules resume <jobId>
```

The open `auto_disabled` incident resolves once a scheduled run succeeds, and the schedule gets a fresh silence window from the moment it is resumed. A code job has no such limit: it is never turned off for failing.

### Limits

* Agent templates do not carry `health` yet, so a template install's schedules are unwatched. The script form of a workflow does not carry it either.
* For a code job, the failure streak is read from its most recent executions (the threshold plus 10).
* Removing a schedule entirely, rather than publishing a version without `health`, leaves an open card in the inbox until the card expires after 30 days.

## If it isn't working

<AccordionGroup>
  <Accordion title="Health shows — although health is declared">
    **Cause** The project was compiled with a lua-cli that does not know `health`; the field was dropped. **Fix** Update lua-cli, compile, and push again. The compiled manifest carries `health` on the schedule when it was kept.
  </Accordion>

  <Accordion title="Health stays unwatched">
    **Cause** The schedule is paused, or the first check has not run yet. **Fix** Resume the schedule, or wait about 5 minutes after publishing.
  </Accordion>

  <Accordion title="An alert every time the schedule runs slowly">
    **Cause** `maxSilenceMinutes` is too close to the cron period plus the run time. **Fix** Raise it with headroom above the longest normal gap.
  </Accordion>
</AccordionGroup>

## Next steps

<Columns cols={2}>
  <Card title="Schedule a workflow" href="/build/workflows/goals-and-schedules">Cadences, schedule input and goals.</Card>
  <Card title="Schedule a recurring job" href="/build/schedule-a-job">A `LuaJob` on a cron schedule.</Card>
  <Card title="Log drains" href="/drains/overview">Send log lines, health lines included, to your own stack.</Card>
  <Card title="Workflows CLI" href="/reference/cli/workflows">`schedules list`, `pause` and `resume`.</Card>
</Columns>
