Skip to main content
After this guide, a scheduled workflow or job says how it is supposed to behave, and the platform tells its owner when it stops: when no scheduled run has succeeded for too long, when runs keep failing, or when the platform has turned the schedule off. You do not need a heartbeat service, a cron monitor or an alert of your own. Compiling health needs the next lua-cli release. A lua-cli that does not know the field drops it without an error, and the schedule is pushed unwatched. Before you begin
1

Declare health

health sits beside schedule, on createWorkflow or on LuaJob.
src/workflows/heartbeat.ts
src/jobs/NightlySyncJob.ts
Give at least one of the two thresholds. A schedule the platform turns off after repeated failures is always an incident while it stays off, which is why a workflow’s failure threshold stops at 3: the platform turns a workflow schedule off after three failed runs in a row, so a fourth never happens.A cron cadence is not parsed, so leave headroom above the cron period plus the run time: */10 with a 5-minute run wants 30, not 15.
2

Compile and push

lua compile checks the declaration and lua push checks it again. Both refuse, with the code shown:The declaration takes effect when the version is live: for a workflow when its version is published (lua push and then activate), for a job when its version is deployed or promoted. Publishing a version without health stops the watch, and an alert that was open for it is closed.
3

Know what happens

The platform checks each schedule that declares health every few minutes.
  1. An incident opens. The owner receives one notice: the schedule’s creator, or the organization’s admins for a schedule nobody owns. It is an inbox card titled with the schedule’s name, plus push or email as notify says. It is a card of its own, separate from the schedule’s run and auto-disable cards.
  2. Later checks stay quiet. Nobody is pinged again while the same incident is open. If the condition changes, for example from failures to the platform turning the schedule off, the open incident is updated in place. A notice that could not be delivered is sent again on the next check and still counts as the same notice.
  3. The incident resolves when the schedule is healthy again: a scheduled run completes inside the window and the failure streak is broken. The same card closes (“Schedule healthy again”) and a recovery notice goes out on the same channels. Pausing the schedule also closes it (“Schedule paused — health alert closed”). Resuming gives the schedule a full window before it can be called silent again.
What counts:
  • Workflows. Only runs fired by the schedule. A manual lua workflows start is neither a success nor a failure. Failed, timed-out and abandoned runs, and fires that could not start, are failures; a cancelled run counts as neither.
  • Jobs. Every execution of the job, including lua jobs trigger. failed, timeout, killed, abandoned and reaped are failures; cancelled and still-running executions are skipped.
4

See it

lua workflows schedules list has a Health column and lua jobs view a Health: line: — when nothing is declared, ok, unwatched, or ALERT <condition> since <time>. With --json, each schedule carries health: { config, state, condition?, since?, lastSucceededAt?, checkedAt? } and lastSucceededAt. state is ok when the schedule is watched and healthy, alerting while an incident is open, and unwatched while the schedule is paused or before its first check, up to about 5 minutes after publishing. condition is silence, failures or auto_disabled.
5

Route alerts to your own tools

Every open, change and resolve is also a line in the agent’s logs: lua logs --type workflow, or --type job for a job. A warn line marks an incident and an info line its resolution. Its fields are health_event, health_condition, job_id, consecutive_failures, last_succeeded_at, silent_for_minutes, max_silence_minutes and alert_after_consecutive_failures.
A log drain carries those lines like any other, so Datadog, Better Stack or your SIEM can alert on health_event opened for every agent in the organization with no per-agent setup.

Options you may need

Resume a schedule the platform turned off

After three failed runs in a row the platform turns a workflow schedule off. Reactivating the workflow does not turn it back on. Fix the cause, then resume the schedule by its job id from lua workflows schedules list:
The open auto_disabled incident resolves once a scheduled run succeeds, and the schedule gets a fresh silence window from the moment it is resumed. A code job has no such limit: it is never turned off for failing.

Limits

  • Agent templates do not carry health yet, so a template install’s schedules are unwatched. The script form of a workflow does not carry it either.
  • For a code job, the failure streak is read from its most recent executions (the threshold plus 10).
  • Removing a schedule entirely, rather than publishing a version without health, leaves an open card in the inbox until the card expires after 30 days.

If it isn’t working

Cause The project was compiled with a lua-cli that does not know health; the field was dropped. Fix Update lua-cli, compile, and push again. The compiled manifest carries health on the schedule when it was kept.
Cause The schedule is paused, or the first check has not run yet. Fix Resume the schedule, or wait about 5 minutes after publishing.
Cause maxSilenceMinutes is too close to the cron period plus the run time. Fix Raise it with headroom above the longest normal gap.

Next steps

Schedule a workflow

Cadences, schedule input and goals.

Schedule a recurring job

A LuaJob on a cron schedule.

Log drains

Send log lines, health lines included, to your own stack.

Workflows CLI

schedules list, pause and resume.