Skip to content
Docs/On-call & escalation

Incidents & alerts

On-call & escalation

Someone should own every outage. Acknowledging an incident stops the reminders. On the Business plan, rotations decide who gets paged and escalation policies make sure someone answers.

Acknowledging incidents

Acknowledging works on every plan. It says “I'm on it”:

  • monitor.reminder “still down” alerts stop for that incident.
  • Escalation stops: no further levels are paged.
  • An internal timeline entry is added: “Acknowledged by Dana: On it, rolling back the last deploy. Reminders and escalation are paused.”
  • incident.acknowledged is sent to the channels that were alerted about the incident: the monitor's channels plus every channel an escalation step paged so far.

Monitor checks, incident updates and subscriber notifications carry on as normal. Resolution still happens automatically when the service recovers.

curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/ack \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"note": "On it, rolling back the last deploy"}'

# A monitor id works too: it acknowledges the monitor's open incident
curl -X POST https://upbutler.com/api/v1/incidents/mon_0n3q8kz1m4hx7c2v9rt/ack \
  -H "Authorization: Bearer $UPBUTLER_API_KEY"

# Changed your mind? Reminders and escalation resume where they stopped
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/unack \
  -H "Authorization: Bearer $UPBUTLER_API_KEY"

The id may be an incident id (inc_…) or a monitor id (mon_…), which acknowledges that monitor's open incident. Acknowledging twice is harmless: the response has alreadyAcknowledged: true. Resolved incidents and maintenance windows can't be acknowledged (409 conflict). Over MCP, use incidents_ack and incidents_unack.

Unacknowledging emits incident.unacknowledged (event stream only) and lets reminders and escalation pick up where they stopped. If a level was due in the meantime, it is paged right away.

Email, Slack, Discord and Telegram alerts for monitor.down, monitor.reminder, incident.created and incident.escalated carry an Acknowledge link while the incident is open and unacknowledged. Links are signed, bound to the incident and the recipient, and valid for 3 days. A signed-in member acknowledges as themselves with one click. Anyone else gets a confirm button (so link scanners can't acknowledge for you), and the timeline credits the recipient, such as “Someone on #ops (slack)”.

On-call schedules

A schedule is a weekly rotation: participants take turns, one week each, handing off on a fixed weekday and time in the schedule's timezone.

curl -X POST https://upbutler.com/api/v1/oncall/schedules \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Primary",
    "timezone": "Europe/Berlin",
    "handoffDay": 0,
    "handoffTime": "09:00",
    "participants": ["usr_0n3q8jy7w2e4r6t8y0u", "usr_0n3q8k01d5f7h9j1k3l", "usr_0n3q8k02m4n6p8q0r2s"],
    "startDay": "2026-10-12"
  }'
FieldTypeDescription
namerequiredstringmax length 80
timezonerequiredstringIANA timezone, e.g. "Europe/Berlin"max length 64
handoffDayrequiredintegerHandoff weekday: 0 = Monday … 6 = Sundaymin 0 · max 6
handoffTimerequiredstringLocal handoff time "HH:MM" (24h)pattern ^\d{1,2}:\d{2}$
participantsrequiredstring[]Ordered member user ids (from GET /members). Each is on call for one week, in this order.max 50 items
startDaystringLocal date (YYYY-MM-DD) whose rotation week the first participant takes. Snapped back to the handoff weekday. Default: the current week.pattern ^\d{4}-\d{2}-\d{2}$
  • Timezone and DST: handoffs happen at the same local wall-clock time all year, so 09:00 Europe/Berlin stays 09:00 across daylight-saving changes. Invalid IANA names return 422.
  • Participants must be members of the workspace (up to 50). An empty list means nobody is on call.
  • Changing participants or handoffTime keeps the rotation's anchor week. Changing handoffDay or sending startDay re-anchors it.
  • Up to 50 schedules per workspace. A schedule used by an escalation policy can't be deleted (409 conflict). Remove it from the policy first.

Who is on call

curl https://upbutler.com/api/v1/oncall/current -H "Authorization: Bearer $UPBUTLER_API_KEY"
200 OK
{
  "data": [
    {
      "scheduleId": "sch_0n3q8k10a2b4c6d8e0f",
      "schedule": "Primary",
      "timezone": "Europe/Berlin",
      "onCall": {
        "userId": "usr_0n3q8k01d5f7h9j1k3l",
        "name": "Dana Weber",
        "email": "[email protected]",
        "until": "2026-10-19T07:00:00.000Z"
      }
    }
  ],
  "at": "2026-10-14T13:02:11.512Z"
}

until is the next moment the person may change: the next handoff, or the start or end of an override. GET /oncall/schedules?days=14 and GET /oncall/schedules/:id add shifts: the upcoming shifts with overrides spliced in, up to 42 days ahead. Agents can use oncall_current to decide whom to tell about a problem.

Overrides

curl -X POST https://upbutler.com/api/v1/oncall/schedules/sch_0n3q8k10a2b4c6d8e0f/overrides \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "userId": "usr_0n3q8jy7w2e4r6t8y0u",
    "start": "2026-10-17T16:00:00+02:00",
    "end": "2026-10-19T09:00:00+02:00",
    "note": "Covering for Dana over the weekend"
  }'

An override puts someone on call for a period, up to 90 days, and wins over the rotation. When overrides overlap, the most recently created one wins. Up to 100 overrides per schedule. Overrides that ended more than 30 days ago are cleaned up automatically. Remove one with DELETE /oncall/schedules/:id/overrides/:overrideId.

Paging contacts

When an escalation level pages a person, either directly or because they're on call, UpButler sends them an email and, if they've set a Telegram chat id, a Telegram message. Each member manages their own contact profile in the dashboard (GET/PATCH /me/contact, session only):

FieldUsed for
emailPaging email. null falls back to the account email
telegramChatIdNumeric Telegram chat id for direct pages
phoneE.164 number such as +491701234567. Stored for upcoming SMS paging, not used yet

GET /oncall/members lists members with their role and whether Telegram and a phone are set up. Paging emails count toward the monthly email quota.

Escalation policies

A policy is an ordered list of levels. Level 1 is paged as soon as an incident opens. Each next level is paged afterMinutes after the previous one, until someone acknowledges or the incident resolves.

curl -X POST https://upbutler.com/api/v1/escalation-policies \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Default",
    "levels": [
      { "afterMinutes": 0,  "targets": [{ "type": "schedule", "id": "sch_0n3q8k10a2b4c6d8e0f" }] },
      { "afterMinutes": 15, "targets": [{ "type": "user", "id": "usr_0n3q8jy7w2e4r6t8y0u" }, { "type": "channel", "id": "ch_0n3q8kz3j6d0q2b5ny8" }] }
    ],
    "repeat": 1,
    "repeatAfterMinutes": 30
  }'

Timeline for this policy, if nobody acknowledges: 0 min, whoever is on call on “Primary” is paged. 15 min, a named person plus a Slack channel. 45 min (15 + 30), the whole chain runs again once.

FieldTypeDescription
namerequiredstringmax length 80
isDefaultbooleanUse for incidents no monitor/page-specific policy claims (the first policy becomes default)
monitorIdsstring[]max 250 items
pageIdsstring[]max 20 items
levelsrequiredobject[]max 10 items
levels[].afterMinutesrequiredintegerWait after the previous level (ignored for level 1)min 0 · max 1,440
levels[].targetsrequiredobject[]max 20 items
levels[].targets[].typerequired"channel"
levels[].targets[].idrequiredstringAlert channel id
repeatintegerRepeat the whole chain this many times while unacknowledgedmin 0 · max 5
repeatAfterMinutesintegermin 1 · max 1,440

Targets

typePages
scheduleWhoever is on call on that schedule at the moment the level fires, by email and Telegram
userA specific member, by email and Telegram
channelAn alert channel (Slack, Discord, email list, webhook, …)

Each level needs 1–20 targets, and a policy has 1–10 levels. repeat (0–5) re-runs the whole chain while unacknowledged, repeatAfterMinutes (default 15) after the last level. On the very first page, channels that already received the monitor.down alert are skipped, so nobody gets the same alert twice. Every step emits incident.escalated with level, round, unacknowledgedForMinutes and who was notified.

Which policy applies

  1. A policy whose monitorIds include one of the incident's monitors.
  2. Otherwise, a policy whose pageIds include one of the incident's status pages.
  3. Otherwise, the default policy (isDefault: true). The first policy you create becomes the default, and marking another one default unsets the previous one.

Escalation applies to incidents opened by monitors, components, agents and the API. Incidents a person declares in the dashboard aren't escalated, because that person is already on it. Incidents that were already more than an hour old when a policy first saw them (for example, a policy created mid-outage) aren't escalated either. Without any policy, alerts follow each monitor's channels and reminders, as before.