health needs the next lua-cli release. A lua-cli that does not know the field drops it without an error, and the schedule is pushed unwatched.
Before you begin
- A workflow with a
schedule(Schedule a workflow and set goals) or aLuaJobwith a recurringschedule(Schedule a recurring job).
1
Declare health
health sits beside schedule, on createWorkflow or on LuaJob.src/workflows/heartbeat.ts
src/jobs/NightlySyncJob.ts
Give at least one of the two thresholds. A schedule the platform turns off after repeated failures is always an incident while it stays off, which is why a workflow’s failure threshold stops at 3: the platform turns a workflow schedule off after three failed runs in a row, so a fourth never happens.A cron cadence is not parsed, so leave headroom above the cron period plus the run time:
*/10 with a 5-minute run wants 30, not 15.2
Compile and push
lua compile checks the declaration and lua push checks it again. Both refuse, with the code shown:The declaration takes effect when the version is live: for a workflow when its version is published (
lua push and then activate), for a job when its version is deployed or promoted. Publishing a version without health stops the watch, and an alert that was open for it is closed.3
Know what happens
The platform checks each schedule that declares
health every few minutes.- An incident opens. The owner receives one notice: the schedule’s creator, or the organization’s admins for a schedule nobody owns. It is an inbox card titled with the schedule’s name, plus push or email as
notifysays. It is a card of its own, separate from the schedule’s run and auto-disable cards. - Later checks stay quiet. Nobody is pinged again while the same incident is open. If the condition changes, for example from failures to the platform turning the schedule off, the open incident is updated in place. A notice that could not be delivered is sent again on the next check and still counts as the same notice.
- The incident resolves when the schedule is healthy again: a scheduled run completes inside the window and the failure streak is broken. The same card closes (“Schedule healthy again”) and a recovery notice goes out on the same channels. Pausing the schedule also closes it (“Schedule paused — health alert closed”). Resuming gives the schedule a full window before it can be called silent again.
- Workflows. Only runs fired by the schedule. A manual
lua workflows startis neither a success nor a failure. Failed, timed-out and abandoned runs, and fires that could not start, are failures; a cancelled run counts as neither. - Jobs. Every execution of the job, including
lua jobs trigger.failed,timeout,killed,abandonedandreapedare failures; cancelled and still-running executions are skipped.
4
See it
lua workflows schedules list has a Health column and lua jobs view a Health: line: — when nothing is declared, ok, unwatched, or ALERT <condition> since <time>. With --json, each schedule carries health: { config, state, condition?, since?, lastSucceededAt?, checkedAt? } and lastSucceededAt. state is ok when the schedule is watched and healthy, alerting while an incident is open, and unwatched while the schedule is paused or before its first check, up to about 5 minutes after publishing. condition is silence, failures or auto_disabled.5
Route alerts to your own tools
Every open, change and resolve is also a line in the agent’s logs: A log drain carries those lines like any other, so Datadog, Better Stack or your SIEM can alert on
lua logs --type workflow, or --type job for a job. A warn line marks an incident and an info line its resolution. Its fields are health_event, health_condition, job_id, consecutive_failures, last_succeeded_at, silent_for_minutes, max_silence_minutes and alert_after_consecutive_failures.health_event opened for every agent in the organization with no per-agent setup.Options you may need
Resume a schedule the platform turned off
After three failed runs in a row the platform turns a workflow schedule off. Reactivating the workflow does not turn it back on. Fix the cause, then resume the schedule by its job id fromlua workflows schedules list:
auto_disabled incident resolves once a scheduled run succeeds, and the schedule gets a fresh silence window from the moment it is resumed. A code job has no such limit: it is never turned off for failing.
Limits
- Agent templates do not carry
healthyet, so a template install’s schedules are unwatched. The script form of a workflow does not carry it either. - For a code job, the failure streak is read from its most recent executions (the threshold plus 10).
- Removing a schedule entirely, rather than publishing a version without
health, leaves an open card in the inbox until the card expires after 30 days.
If it isn’t working
Health shows — although health is declared
Health shows — although health is declared
Cause The project was compiled with a lua-cli that does not know
health; the field was dropped. Fix Update lua-cli, compile, and push again. The compiled manifest carries health on the schedule when it was kept.Health stays unwatched
Health stays unwatched
Cause The schedule is paused, or the first check has not run yet. Fix Resume the schedule, or wait about 5 minutes after publishing.
An alert every time the schedule runs slowly
An alert every time the schedule runs slowly
Cause
maxSilenceMinutes is too close to the cron period plus the run time. Fix Raise it with headroom above the longest normal gap.Next steps
Schedule a workflow
Cadences, schedule input and goals.
Schedule a recurring job
A
LuaJob on a cron schedule.Log drains
Send log lines, health lines included, to your own stack.
Workflows CLI
schedules list, pause and resume.
