Skip to content

Cron job monitoring that knows when a job didn't run

A cron job that stops does not throw an error. It just stops. Heartbeat monitoring turns that silence into an alert: your job pings a URL when it finishes, and when the ping does not come, or the job says it failed, you hear about it.

A dead man's switch in two lines

Create a monitor with a periodSec and you get a ping URL back. Append one curl to the job. Set periodSec to the cron period (86,400 for daily) and graceSec to the job's normal runtime plus some slack.

The monitor goes down when no ping arrives within periodSec + graceSec of the previous one. A missed deadline goes through the normal confirmation; an explicit /fail ping skips it and alerts immediately. The next success ping recovers it.

periodSec
300
How often the job is expected to ping: 30 seconds to 31 days
graceSec
60
Extra slack before a missing ping counts as down: 0 to 86,400 seconds
curl -X POST https://upbutler.com/api/v1/monitors \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "Nightly backup", "periodSec": 86400, "graceSec": 1800}'

Wherever the job runs

Anything that can make an HTTPS request when it finishes: crontab, CI schedules, serverless cron routes, queue workers, backups and the loop of a long-running agent. No SDK and no API key in the job.

name: nightly-sync
on:
  schedule: [{ cron: "0 3 * * *" }]
jobs:
  sync:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: ./scripts/sync.sh
      - name: Heartbeat (success)
        if: success()
        run: curl -fsS -m 10 --retry 3 "${{ secrets.UPBUTLER_HB_URL }}?msg=run+${{ github.run_id }}"
      - name: Heartbeat (failure)
        if: failure()
        run: curl -fsS -m 10 --retry 3 "${{ secrets.UPBUTLER_HB_URL }}/fail?msg=run+${{ github.run_id }}+failed"

Or let it be found for you. npx upbutler init scans a repository, creates a heartbeat for each cron job it finds and writes upbutler.yaml so they stay in sync. An agent can do the same with one call to POST /api/v1/services, which returns ready-made pingUrl and failUrl.

More than "it ran"

GET or POST both work and both return JSON. What the job tells the ping URL is stored with the check and shown in the dashboard, in alerts and in the incident.

RequestReports
/hb/<token>Success
/hb/<token>/failFailure (down). Alerts immediately, with no confirmation delay
?status=down · ?status=degradedFailure, or "it ran, but not well"
?msg=…A message stored with the ping (up to 500 characters)
POST with a plain-text bodyThe body becomes the message
POST /api/v1/heartbeat/<token>JSON with status, message and durationMs; the duration is charted as the job's runtime

From there a heartbeat is a monitor like any other: the same alert channels, the same incidents with an AI-written report, and a status page component if you want customers to see that the nightly sync is late. When the scheduler itself is the problem, the service status catalog shows it: Vercel Cron Jobs, GitHub Actions, Render cron jobs.

Limits, plainly

Free$0

10

monitors of any kind, heartbeats included

Starter$12/mo

25

monitors of any kind, heartbeats included

Pro$29/mo

75

monitors of any kind, heartbeats included

Business$99/mo

300

monitors of any kind, heartbeats included

A heartbeat counts as one monitor. New workspaces start with 14 days of Pro. See pricing, or uptime monitoring for the checks that reach your service from outside.

Cron monitoring questions

What is cron job monitoring?

Cron job monitoring tells you when a scheduled job did not run, ran late or failed. The job cannot be checked from outside, so it works the other way round: the job requests a unique URL each time it finishes, and the monitor alerts you when that request does not arrive in time. This is also called heartbeat monitoring or a dead man's switch.

When exactly is a job considered down?

When no ping arrives within periodSec + graceSec of the previous one (or of the monitor's creation, for the first ping). periodSec is how often the job should run (30 seconds to 31 days, 300 by default) and graceSec is extra slack for its runtime (up to 86,400 seconds, 60 by default).

What if the job runs but fails?

Have it request the /fail URL, or send ?status=down. An explicit failure skips the usual confirmation and alerts immediately, with your message attached. ?status=degraded reports a run that finished but not cleanly. A success ping after an outage recovers the monitor immediately.

Do I need an API key or an SDK in the job?

No. Ping URLs need no API key and work with a plain curl, GET or POST. Treat the hb_… token like a password: keep it in CI secrets or environment variables. Each token accepts up to 120 pings per minute.

Does it work with Vercel Cron, GitHub Actions and Kubernetes CronJobs?

Yes: anything that can make an HTTPS request when it finishes. Add a fetch at the end of a Vercel Cron route, a curl step to a GitHub Actions workflow, or a curl after the command of a Kubernetes CronJob. npx upbutler init finds the cron routes in a repository and creates the heartbeats for you.

Can my customers see whether a job is running?

If you want them to. Add a status page component that follows the heartbeat monitor, or pass statusPage when you create it, and the component turns red when the job misses its deadline. Leave it off the page and it stays internal.

Can I record how long the job took?

Yes. Post JSON to /api/v1/heartbeat/<token> with durationMs; it is stored as the check's latency, so job durations show up in the same charts as response times.

Is cron job monitoring free?

Heartbeats are ordinary monitors and are included on every plan: 10 monitors on Free, 25 on Starter, 75 on Pro and 300 on Business. Commercial use is fine on the free plan.

One curl at the end of the job.

Free for 10 monitors, heartbeats included, with a status page. No card.