Skip to content
Docs/Heartbeats

Monitoring

Heartbeats

For things we can't reach from outside: cron jobs, queue workers, CI pipelines, backups and agent loops. Your job pings UpButler after it runs. If the pings stop, or the job reports a failure, you get alerted.

Create a heartbeat

curl -X POST https://upbutler.com/api/v1/monitors \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "Nightly backup", "periodSec": 86400, "graceSec": 1800}'

Any monitor with periodSec is a heartbeat. The response includes heartbeat.pingUrl. POST /services also returns ready-made pingUrl, failUrl and a curl example.

FieldDefaultRangeMeaning
periodSec30030 s – 31 daysHow often your job is expected to ping
graceSec600 – 86,400Extra slack before a missing ping counts as down

The monitor goes down when no ping arrives within periodSec + graceSec of the previous one, or of the monitor's creation for the first ping. Missed deadlines go through the normal confirmation (failureThreshold). Explicit failure pings skip it and alert immediately. A success ping after an outage recovers immediately.

Ping URL reference

GET or POST both work, and both return JSON.

RequestReports
/hb/<token>Success
/hb/<token>/failFailure (down), alerts immediately
?status=down · ?status=degradedFailure / degraded
?msg=… (or ?message=…)A message stored with the ping (max 500 chars)
POST plain-text bodyThe body becomes the message (first 2,000 bytes are read)
POST JSON bodymessage, and status: down/fail/failed → down, degraded → degraded
Short URLs
# Success
curl -fsS https://upbutler.com/hb/hb_Xf3kq9LmR2vT8wYzN4bC1dE7

# Success with a message (stored with the ping, shown in the dashboard)
curl -fsS "https://upbutler.com/hb/hb_Xf3kq9LmR2vT8wYzN4bC1dE7?msg=backup+42GB+ok"

# Explicit failure: alerts immediately, no confirmation delay
curl -fsS https://upbutler.com/hb/hb_Xf3kq9LmR2vT8wYzN4bC1dE7/fail
curl -fsS "https://upbutler.com/hb/hb_Xf3kq9LmR2vT8wYzN4bC1dE7?status=down&msg=pg_dump+exit+code+1"

# Degraded (it ran, but not well)
curl -fsS "https://upbutler.com/hb/hb_Xf3kq9LmR2vT8wYzN4bC1dE7?status=degraded&msg=3+of+120+files+skipped"

# POST a plain-text body: it becomes the message
curl -fsS -X POST https://upbutler.com/hb/hb_Xf3kq9LmR2vT8wYzN4bC1dE7 --data-binary "rotated 1,204 log files"

# POST JSON: {"status": "down" | "fail" | "failed" | "degraded", "message": "..."}
curl -fsS -X POST https://upbutler.com/hb/hb_Xf3kq9LmR2vT8wYzN4bC1dE7 -H "Content-Type: application/json" \
  -d '{"status": "fail", "message": "S3 upload timed out"}'

The response is {"ok": true, "monitor": "Nightly backup", "status": "up", "receivedAt": "…"}. An unknown token returns 404 not_found. A paused monitor accepts the ping but doesn't change state.

API form, with durations

POST /api/v1/heartbeat/:token (MCP tool heartbeat_ping) takes JSON with status (up | down | degraded), message and durationMs, which is stored as the check's latency so job durations show up in charts. It also needs no API key.

curl -X POST https://upbutler.com/api/v1/heartbeat/hb_Xf3kq9LmR2vT8wYzN4bC1dE7 \
  -H "Content-Type: application/json" \
  -d '{"status": "up", "message": "indexed 18,402 docs", "durationMs": 48210}'
200 OK
{
  "ok": true,
  "monitor": "Nightly backup",
  "receivedAt": "2026-10-09T02:00:51.204Z",
  "via": "rest"
}

Recipes

crontab

# Ping only when the job succeeds, and report a failure otherwise
0 2 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 3 https://upbutler.com/hb/hb_Xf3kq9LmR2vT8wYzN4bC1dE7 >/dev/null \
  || curl -fsS -m 10 --retry 3 https://upbutler.com/hb/hb_Xf3kq9LmR2vT8wYzN4bC1dE7/fail >/dev/null

Set periodSec to the cron period (86,400 for daily) and graceSec to the job's normal runtime plus some slack.

Wrap any command (with duration and output)

hb-run
#!/usr/bin/env bash
# hb-run: run any command and report the result + duration to UpButler.
#   hb-run hb_Xf3kq9LmR2vT8wYzN4bC1dE7 ./backup.sh --full
set -uo pipefail
token="$1"; shift
start=$(date +%s%3N)
output=$("$@" 2>&1); code=$?
duration=$(( $(date +%s%3N) - start ))
status=$([ $code -eq 0 ] && echo up || echo down)
curl -fsS -m 10 --retry 3 -X POST "https://upbutler.com/api/v1/heartbeat/$token" \
  -H "Content-Type: application/json" \
  -d "$(jq -n --arg s "$status" --arg m "${output: -400}" --argjson d "$duration" \
        '{status: $s, message: $m, durationMs: $d}')" >/dev/null
exit $code

GitHub Actions

.github/workflows/nightly-sync.yml
name: nightly-sync
on:
  schedule: [{ cron: "0 3 * * *" }]
jobs:
  sync:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: ./scripts/sync.sh
      - name: Heartbeat (success)
        if: success()
        run: curl -fsS -m 10 --retry 3 "${{ secrets.UPBUTLER_HB_URL }}?msg=run+${{ github.run_id }}"
      - name: Heartbeat (failure)
        if: failure()
        run: curl -fsS -m 10 --retry 3 "${{ secrets.UPBUTLER_HB_URL }}/fail?msg=run+${{ github.run_id }}+failed"

Agent loops and workers

import time, requests

HB = "https://upbutler.com/hb/hb_Xf3kq9LmR2vT8wYzN4bC1dE7"

def tick():
    ...  # your agent's unit of work

while True:
    started = time.monotonic()
    try:
        tick()
        requests.post(HB, json={"message": "tick ok"}, timeout=10)
    except Exception as e:
        # Explicit failure: alerts right away instead of waiting for the deadline
        requests.post(f"{HB}/fail", data=str(e)[:500], timeout=10)
    time.sleep(max(0, 300 - (time.monotonic() - started)))   # periodSec = 300

Long-running workers should ping on every iteration (or every N minutes), with periodSec slightly above the loop interval. To show the job on a status page, add a component that follows the heartbeat monitor, or pass statusPage to POST /services.