Incidents & alerts
On-call & escalation
Someone should own every outage. Acknowledging an incident stops the reminders. On the Business plan, rotations decide who gets paged and escalation policies make sure someone answers.
Acknowledging incidents
Acknowledging works on every plan. It says “I'm on it”:
monitor.reminder“still down” alerts stop for that incident.- Escalation stops: no further levels are paged.
- An internal timeline entry is added: “Acknowledged by Dana: On it, rolling back the last deploy. Reminders and escalation are paused.”
incident.acknowledgedis sent to the channels that were alerted about the incident: the monitor's channels plus every channel an escalation step paged so far.
Monitor checks, incident updates and subscriber notifications carry on as normal. Resolution still happens automatically when the service recovers.
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/ack \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"note": "On it, rolling back the last deploy"}'
# A monitor id works too: it acknowledges the monitor's open incident
curl -X POST https://upbutler.com/api/v1/incidents/mon_0n3q8kz1m4hx7c2v9rt/ack \
-H "Authorization: Bearer $UPBUTLER_API_KEY"
# Changed your mind? Reminders and escalation resume where they stopped
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/unack \
-H "Authorization: Bearer $UPBUTLER_API_KEY"The id may be an incident id (inc_…) or a monitor id (mon_…), which acknowledges that monitor's open incident. Acknowledging twice is harmless: the response has alreadyAcknowledged: true. Resolved incidents and maintenance windows can't be acknowledged (409 conflict). Over MCP, use incidents_ack and incidents_unack.
Unacknowledging emits incident.unacknowledged (event stream only) and lets reminders and escalation pick up where they stopped. If a level was due in the meantime, it is paged right away.
Acknowledge links in alerts
Email, Slack, Discord and Telegram alerts for monitor.down, monitor.reminder, incident.created and incident.escalated carry an Acknowledge link while the incident is open and unacknowledged. Links are signed, bound to the incident and the recipient, and valid for 3 days. A signed-in member acknowledges as themselves with one click. Anyone else gets a confirm button (so link scanners can't acknowledge for you), and the timeline credits the recipient, such as “Someone on #ops (slack)”.
On-call schedules
A schedule is a weekly rotation: participants take turns, one week each, handing off on a fixed weekday and time in the schedule's timezone.
curl -X POST https://upbutler.com/api/v1/oncall/schedules \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Primary",
"timezone": "Europe/Berlin",
"handoffDay": 0,
"handoffTime": "09:00",
"participants": ["usr_0n3q8jy7w2e4r6t8y0u", "usr_0n3q8k01d5f7h9j1k3l", "usr_0n3q8k02m4n6p8q0r2s"],
"startDay": "2026-10-12"
}'| Field | Type | Description |
|---|---|---|
namerequired | string | max length 80 |
timezonerequired | string | IANA timezone, e.g. "Europe/Berlin"max length 64 |
handoffDayrequired | integer | Handoff weekday: 0 = Monday … 6 = Sundaymin 0 · max 6 |
handoffTimerequired | string | Local handoff time "HH:MM" (24h)pattern ^\d{1,2}:\d{2}$ |
participantsrequired | string[] | Ordered member user ids (from GET /members). Each is on call for one week, in this order.max 50 items |
startDay | string | Local date (YYYY-MM-DD) whose rotation week the first participant takes. Snapped back to the handoff weekday. Default: the current week.pattern ^\d{4}-\d{2}-\d{2}$ |
- Timezone and DST: handoffs happen at the same local wall-clock time all year, so
09:00 Europe/Berlinstays 09:00 across daylight-saving changes. Invalid IANA names return422. - Participants must be members of the workspace (up to 50). An empty list means nobody is on call.
- Changing
participantsorhandoffTimekeeps the rotation's anchor week. ChanginghandoffDayor sendingstartDayre-anchors it. - Up to 50 schedules per workspace. A schedule used by an escalation policy can't be deleted (
409 conflict). Remove it from the policy first.
Who is on call
curl https://upbutler.com/api/v1/oncall/current -H "Authorization: Bearer $UPBUTLER_API_KEY"{
"data": [
{
"scheduleId": "sch_0n3q8k10a2b4c6d8e0f",
"schedule": "Primary",
"timezone": "Europe/Berlin",
"onCall": {
"userId": "usr_0n3q8k01d5f7h9j1k3l",
"name": "Dana Weber",
"email": "[email protected]",
"until": "2026-10-19T07:00:00.000Z"
}
}
],
"at": "2026-10-14T13:02:11.512Z"
}until is the next moment the person may change: the next handoff, or the start or end of an override. GET /oncall/schedules?days=14 and GET /oncall/schedules/:id add shifts: the upcoming shifts with overrides spliced in, up to 42 days ahead. Agents can use oncall_current to decide whom to tell about a problem.
Overrides
curl -X POST https://upbutler.com/api/v1/oncall/schedules/sch_0n3q8k10a2b4c6d8e0f/overrides \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"userId": "usr_0n3q8jy7w2e4r6t8y0u",
"start": "2026-10-17T16:00:00+02:00",
"end": "2026-10-19T09:00:00+02:00",
"note": "Covering for Dana over the weekend"
}'An override puts someone on call for a period, up to 90 days, and wins over the rotation. When overrides overlap, the most recently created one wins. Up to 100 overrides per schedule. Overrides that ended more than 30 days ago are cleaned up automatically. Remove one with DELETE /oncall/schedules/:id/overrides/:overrideId.
Paging contacts
When an escalation level pages a person, either directly or because they're on call, UpButler sends them an email and, if they've set a Telegram chat id, a Telegram message. Each member manages their own contact profile in the dashboard (GET/PATCH /me/contact, session only):
| Field | Used for |
|---|---|
email | Paging email. null falls back to the account email |
telegramChatId | Numeric Telegram chat id for direct pages |
phone | E.164 number such as +491701234567. Stored for upcoming SMS paging, not used yet |
GET /oncall/members lists members with their role and whether Telegram and a phone are set up. Paging emails count toward the monthly email quota.
Escalation policies
A policy is an ordered list of levels. Level 1 is paged as soon as an incident opens. Each next level is paged afterMinutes after the previous one, until someone acknowledges or the incident resolves.
curl -X POST https://upbutler.com/api/v1/escalation-policies \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Default",
"levels": [
{ "afterMinutes": 0, "targets": [{ "type": "schedule", "id": "sch_0n3q8k10a2b4c6d8e0f" }] },
{ "afterMinutes": 15, "targets": [{ "type": "user", "id": "usr_0n3q8jy7w2e4r6t8y0u" }, { "type": "channel", "id": "ch_0n3q8kz3j6d0q2b5ny8" }] }
],
"repeat": 1,
"repeatAfterMinutes": 30
}'Timeline for this policy, if nobody acknowledges: 0 min, whoever is on call on “Primary” is paged. 15 min, a named person plus a Slack channel. 45 min (15 + 30), the whole chain runs again once.
| Field | Type | Description |
|---|---|---|
namerequired | string | max length 80 |
isDefault | boolean | Use for incidents no monitor/page-specific policy claims (the first policy becomes default) |
monitorIds | string[] | max 250 items |
pageIds | string[] | max 20 items |
levelsrequired | object[] | max 10 items |
levels[].afterMinutesrequired | integer | Wait after the previous level (ignored for level 1)min 0 · max 1,440 |
levels[].targetsrequired | object[] | max 20 items |
levels[].targets[].typerequired | "channel" | |
levels[].targets[].idrequired | string | Alert channel id |
repeat | integer | Repeat the whole chain this many times while unacknowledgedmin 0 · max 5 |
repeatAfterMinutes | integer | min 1 · max 1,440 |
Targets
type | Pages |
|---|---|
schedule | Whoever is on call on that schedule at the moment the level fires, by email and Telegram |
user | A specific member, by email and Telegram |
channel | An alert channel (Slack, Discord, email list, webhook, …) |
Each level needs 1–20 targets, and a policy has 1–10 levels. repeat (0–5) re-runs the whole chain while unacknowledged, repeatAfterMinutes (default 15) after the last level. On the very first page, channels that already received the monitor.down alert are skipped, so nobody gets the same alert twice. Every step emits incident.escalated with level, round, unacknowledgedForMinutes and who was notified.
Which policy applies
- A policy whose
monitorIdsinclude one of the incident's monitors. - Otherwise, a policy whose
pageIdsinclude one of the incident's status pages. - Otherwise, the default policy (
isDefault: true). The first policy you create becomes the default, and marking another one default unsets the previous one.
Escalation applies to incidents opened by monitors, components, agents and the API. Incidents a person declares in the dashboard aren't escalated, because that person is already on it. Incidents that were already more than an hour old when a policy first saw them (for example, a policy created mid-outage) aren't escalated either. Without any policy, alerts follow each monitor's channels and reminders, as before.