Cron job monitoring that knows when a job didn't run
A cron job that stops does not throw an error. It just stops. Heartbeat monitoring turns that silence into an alert: your job pings a URL when it finishes, and when the ping does not come, or the job says it failed, you hear about it.
| 02:00:00 | backup.sh starts | cron fires as scheduled | |
| 02:00:51 | ping received | "backup 42GB ok" · 51 s | |
| +1 day 02:00 | no ping | the server was rebuilt; the crontab was not | |
| +1 day 02:30 | deadline passed | periodSec + graceSec without a ping: down | |
| +1 day 02:30 | incident opened | Slack, email · status page component follows |
A dead man's switch in two lines
Create a monitor with a periodSec and you get a ping URL back. Append one curl to the job. Set periodSec to the cron period (86,400 for daily) and graceSec to the job's normal runtime plus some slack.
The monitor goes down when no ping arrives within periodSec + graceSec of the previous one. A missed deadline goes through the normal confirmation; an explicit /fail ping skips it and alerts immediately. The next success ping recovers it.
- periodSec
- 300
- How often the job is expected to ping: 30 seconds to 31 days
- graceSec
- 60
- Extra slack before a missing ping counts as down: 0 to 86,400 seconds
curl -X POST https://upbutler.com/api/v1/monitors \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name": "Nightly backup", "periodSec": 86400, "graceSec": 1800}'# Ping only when the job succeeds, and report a failure otherwise
0 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 3 https://upbutler.com/hb/<token> >/dev/null \
|| curl -fsS -m 10 --retry 3 https://upbutler.com/hb/<token>/fail >/dev/nullWherever the job runs
Anything that can make an HTTPS request when it finishes: crontab, CI schedules, serverless cron routes, queue workers, backups and the loop of a long-running agent. No SDK and no API key in the job.
name: nightly-sync
on:
schedule: [{ cron: "0 3 * * *" }]
jobs:
sync:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: ./scripts/sync.sh
- name: Heartbeat (success)
if: success()
run: curl -fsS -m 10 --retry 3 "${{ secrets.UPBUTLER_HB_URL }}?msg=run+${{ github.run_id }}"
- name: Heartbeat (failure)
if: failure()
run: curl -fsS -m 10 --retry 3 "${{ secrets.UPBUTLER_HB_URL }}/fail?msg=run+${{ github.run_id }}+failed"// Node 18+ / Bun / Deno: a Vercel Cron route, a queue worker, a script
const HB = process.env.UPBUTLER_HB_URL!;
try {
await runJob();
await fetch(HB, { method: "POST", body: "ok" });
} catch (err) {
await fetch(`${HB}/fail`, { method: "POST", body: String(err).slice(0, 500) });
throw err;
}import time, requests
HB = "https://upbutler.com/hb/<token>"
while True:
started = time.monotonic()
try:
tick() # your agent's unit of work
requests.post(HB, json={"message": "tick ok"}, timeout=10)
except Exception as e:
# Explicit failure: alerts right away instead of waiting for the deadline
requests.post(f"{HB}/fail", data=str(e)[:500], timeout=10)
time.sleep(max(0, 300 - (time.monotonic() - started))) # periodSec = 300Or let it be found for you. npx upbutler init scans a repository, creates a heartbeat for each cron job it finds and writes upbutler.yaml so they stay in sync. An agent can do the same with one call to POST /api/v1/services, which returns ready-made pingUrl and failUrl.
More than "it ran"
GET or POST both work and both return JSON. What the job tells the ping URL is stored with the check and shown in the dashboard, in alerts and in the incident.
| Request | Reports |
|---|---|
| /hb/<token> | Success |
| /hb/<token>/fail | Failure (down). Alerts immediately, with no confirmation delay |
| ?status=down · ?status=degraded | Failure, or "it ran, but not well" |
| ?msg=… | A message stored with the ping (up to 500 characters) |
| POST with a plain-text body | The body becomes the message |
| POST /api/v1/heartbeat/<token> | JSON with status, message and durationMs; the duration is charted as the job's runtime |
From there a heartbeat is a monitor like any other: the same alert channels, the same incidents with an AI-written report, and a status page component if you want customers to see that the nightly sync is late. When the scheduler itself is the problem, the service status catalog shows it: Vercel Cron Jobs, GitHub Actions, Render cron jobs.
Limits, plainly
Free$0
10
monitors of any kind, heartbeats included
Starter$12/mo
25
monitors of any kind, heartbeats included
Pro$29/mo
75
monitors of any kind, heartbeats included
Business$99/mo
300
monitors of any kind, heartbeats included
A heartbeat counts as one monitor. New workspaces start with 14 days of Pro. See pricing, or uptime monitoring for the checks that reach your service from outside.
Cron monitoring questions
What is cron job monitoring?
Cron job monitoring tells you when a scheduled job did not run, ran late or failed. The job cannot be checked from outside, so it works the other way round: the job requests a unique URL each time it finishes, and the monitor alerts you when that request does not arrive in time. This is also called heartbeat monitoring or a dead man's switch.
When exactly is a job considered down?
When no ping arrives within periodSec + graceSec of the previous one (or of the monitor's creation, for the first ping). periodSec is how often the job should run (30 seconds to 31 days, 300 by default) and graceSec is extra slack for its runtime (up to 86,400 seconds, 60 by default).
What if the job runs but fails?
Have it request the /fail URL, or send ?status=down. An explicit failure skips the usual confirmation and alerts immediately, with your message attached. ?status=degraded reports a run that finished but not cleanly. A success ping after an outage recovers the monitor immediately.
Do I need an API key or an SDK in the job?
No. Ping URLs need no API key and work with a plain curl, GET or POST. Treat the hb_… token like a password: keep it in CI secrets or environment variables. Each token accepts up to 120 pings per minute.
Does it work with Vercel Cron, GitHub Actions and Kubernetes CronJobs?
Yes: anything that can make an HTTPS request when it finishes. Add a fetch at the end of a Vercel Cron route, a curl step to a GitHub Actions workflow, or a curl after the command of a Kubernetes CronJob. npx upbutler init finds the cron routes in a repository and creates the heartbeats for you.
Can my customers see whether a job is running?
If you want them to. Add a status page component that follows the heartbeat monitor, or pass statusPage when you create it, and the component turns red when the job misses its deadline. Leave it off the page and it stays internal.
Can I record how long the job took?
Yes. Post JSON to /api/v1/heartbeat/<token> with durationMs; it is stored as the check's latency, so job durations show up in the same charts as response times.
Is cron job monitoring free?
Heartbeats are ordinary monitors and are included on every plan: 10 monitors on Free, 25 on Starter, 75 on Pro and 300 on Business. Commercial use is fine on the free plan.
One curl at the end of the job.
Free for 10 monitors, heartbeats included, with a status page. No card.